QdrantVFS exposes a Qdrant collection as a read-only filesystem: group-by
payload fields become nested folders, each point is a .json payload file (plus
a .txt text file and an optional blob), and semantic search is the search
command. The TypeScript backend mirrors the Python one
and returns identical results.
Install
The client is browser-safe, so the VFS ships incore and is available
from both the Node and browser packages:
Document and chunk lineage
Config fields accept Qdrant’s dotted nested-key syntax. A LangChain-style payload withpage_content and metadata: { source, page } can therefore use:
s3://docs/policies/refund-2026.pdf and page value 004,
the chunk text is exposed as
refund-2026.pdf/004__<point-id>.txt. The point id remains as a stable suffix,
so duplicate labels cannot collide and direct reads work without a warm cache.
Only the stem the listing publishes opens: another label in front of the same
id reads as absent. A label that is not a string spells as compact JSON
(true, 1, 1e-7), the same in TypeScript and Python. basenameFields strips
URL/path parents from the named groupBy fields; omit a field to preserve its
complete value in one path-safe segment: / renders as ∕, and a blank or
dot-led value is led by ⁄, so every value has its own segment that lists and
opens. A basename longer than 255 bytes, which ext4
and APFS refuse, is cut to fit and ends in __ plus the md5 of the whole name,
so two long leaves stay two directories. Basenames must be unique within a
parent group; ambiguous listings are refused, and opening a basename directory
checks every point of its parent group, so a second source past maxRows is
refused rather than hidden.
Filesystem layout
<id> is the Qdrant point id. With nameField, leaf stems use
<name>__<id>. Semantic search is the search command, not a path: it returns
ranked points as the canonical .txt (or .json) paths
above, annotated with the similarity score.
Search embedding
search needs a vector for the query, and the config decides where it comes
from. With embed set, a (text: string) => Promise<number[]> the caller
brings, the query is vectorized in-process, so a self-hosted Qdrant works and
mirage depends on no model runtime; examples/typescript/qdrant/qdrant_vfs.ts
feeds it a dependency-free hashed bag of words, and any model that returns a
vector plugs in the same way. Without it, cloudInference: true sends the query
text to the server, so the cluster must have inference enabled (Qdrant Cloud)
and store vectors from the same embeddingModel (default
sentence-transformers/all-MiniLM-L6-v2). With neither, search refuses the
query: Python embeds it in process by default, which no JS client can.
The hook never lands in a snapshot: a restored mount asks for a fresh
VFS. Browsing (ls/cat/find/grep) needs no embedding.
Supported commands
ls, cd, tree, cat, stat, find, wc, head, tail, and search. grep/rg stay
lexical; search "<query>" <path> is the semantic command, returning ranked
points as canonical <id>.txt (or <id>.json) paths plus a score, so results
compose with cat, wc, and pipes. Flags: --top-k, --threshold, --method semantic.
Folder listings filter on payload fields. A filtered listing scrolls first and
only creates keyword payload indexes for the groupBy fields if Qdrant reports
one is required. maxRows caps how many points are scanned per folder.