Skip to main content
QdrantVFS exposes a Qdrant collection as a read-only filesystem: group-by payload fields become nested folders, each point is a .json payload file (plus a .txt text file and an optional blob), and semantic search is the search command. The TypeScript backend mirrors the Python one and returns identical results.

Install

The client is browser-safe, so the VFS ships in core and is available from both the Node and browser packages:

Document and chunk lineage

Config fields accept Qdrant’s dotted nested-key syntax. A LangChain-style payload with page_content and metadata: { source, page } can therefore use:
For a source value s3://docs/policies/refund-2026.pdf and page value 004, the chunk text is exposed as refund-2026.pdf/004__<point-id>.txt. The point id remains as a stable suffix, so duplicate labels cannot collide and direct reads work without a warm cache. Only the stem the listing publishes opens: another label in front of the same id reads as absent. A label that is not a string spells as compact JSON (true, 1, 1e-7), the same in TypeScript and Python. basenameFields strips URL/path parents from the named groupBy fields; omit a field to preserve its complete value in one path-safe segment: / renders as ∕, and a blank or dot-led value is led by ⁄, so every value has its own segment that lists and opens. A basename longer than 255 bytes, which ext4 and APFS refuse, is cut to fit and ends in __ plus the md5 of the whole name, so two long leaves stay two directories. Basenames must be unique within a parent group; ambiguous listings are refused, and opening a basename directory checks every point of its parent group, so a second source past maxRows is refused rather than hidden.

Filesystem layout

<id> is the Qdrant point id. With nameField, leaf stems use <name>__<id>. Semantic search is the search command, not a path: it returns ranked points as the canonical .txt (or .json) paths above, annotated with the similarity score.

Search embedding

search needs a vector for the query, and the config decides where it comes from. With embed set, a (text: string) => Promise<number[]> the caller brings, the query is vectorized in-process, so a self-hosted Qdrant works and mirage depends on no model runtime; examples/typescript/qdrant/qdrant_vfs.ts feeds it a dependency-free hashed bag of words, and any model that returns a vector plugs in the same way. Without it, cloudInference: true sends the query text to the server, so the cluster must have inference enabled (Qdrant Cloud) and store vectors from the same embeddingModel (default sentence-transformers/all-MiniLM-L6-v2). With neither, search refuses the query: Python embeds it in process by default, which no JS client can. The hook never lands in a snapshot: a restored mount asks for a fresh VFS. Browsing (ls/cat/find/grep) needs no embedding.

Supported commands

ls, cd, tree, cat, stat, find, wc, head, tail, and search. grep/rg stay lexical; search "<query>" <path> is the semantic command, returning ranked points as canonical <id>.txt (or <id>.json) paths plus a score, so results compose with cat, wc, and pipes. Flags: --top-k, --threshold, --method semantic. Folder listings filter on payload fields. A filtered listing scrolls first and only creates keyword payload indexes for the groupBy fields if Qdrant reports one is required. maxRows caps how many points are scanned per folder.