> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mirage.strukto.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Adding a New VFS

> Ship your own backend on the public `mirage` surface, or contribute a builtin VFS using Mirage's VFS, command, and snapshot conventions.

A VFS maps an external system to Mirage's filesystem operations and shell commands. There are two paths:

* **Ship your own backend** — a single Python file in your own project or package, built on what the `mirage` package exports at its root. No Mirage fork, no edits to Mirage source.
* **Contribute a builtin** — the four-layer layout inside the Mirage repo, mirrored in TypeScript.

## Ship Your Own Backend

Everything an out-of-tree backend needs is exported from the `mirage` package root, the same front door that hands out `Workspace` and the twin of what `@struktoai/mirage-core` gives a TypeScript backend; there is no separate SDK module. Write the core functions over your data source, group the three required reads in a `VFSAdapter`, and `BaseVFS` wires the full generic command set (`ls`, `cat`, `grep`, `find`, `head`, `wc`, ...) plus glob resolution:

```python theme={null}
from pydantic import BaseModel, ConfigDict, SecretStr

from mirage import (Accessor, BaseVFS, FileStat, ReadOps, VFSAdapter)


class JiraConfig(BaseModel):
    model_config = ConfigDict(extra="forbid")

    site: str
    token: SecretStr


class JiraAccessor(Accessor):
    def __init__(self, config: JiraConfig) -> None:
        self.client = make_client(config)

    async def close(self) -> None:
        await self.client.close()


async def readdir(accessor, path, index=None) -> list[str]: ...
async def read_bytes(accessor, path, index=None) -> bytes: ...
async def stat(accessor, path, index=None) -> FileStat: ...


class JiraVFS(BaseVFS):
    CONFIG_CLS = JiraConfig

    def __init__(self, config: JiraConfig) -> None:
        super().__init__(
            name="jira",
            accessor=JiraAccessor(config),
            io=VFSAdapter(read=ReadOps(
                readdir=readdir,
                read_bytes=read_bytes,
                stat=stat,
            )),
            prompt="Issues rendered as .json files.",
        )

```

Mount it like any builtin: `Workspace({"/jira/": JiraVFS(cfg)})`. A class rather than a
factory function, because that is what the registry and the config reference below both name;
`examples/python/other/custom_vfs.py` is the same shape end to end.

For a smaller runnable example, see `examples/python/other/basevfs_views.py`.
Its read-only `NotesVFS` connects to a nested mount, a namespace symlink,
a CLI using namespace and session views, and the filesystem and runtime APIs.

### Add capabilities as the resource grows

Only `readdir`, `read_bytes`, and `stat` are required. The adapter derives a
stream from `read_bytes`, existence checks from `stat`, and defaults to a
remote resource that is available. Range reads fall back to reading and slicing.
A derived stream still fetches the entire file; use a native stream for large files.

```python theme={null}
from mirage import NativeReadOps, ReadOps, VFSAdapter, WriteOps

adapter = VFSAdapter(
    read=ReadOps(readdir=readdir, read_bytes=read_bytes, stat=stat),
    native=NativeReadOps(read_stream=read_stream, read_range=read_range),
    writes=WriteOps(write=write, unlink=unlink),
)
```

Each callback above is implemented by your backend. Leave out `native` and
`writes` for a minimal read-only adapter. These groups are independent:

| Group | Operations | Purpose |
| - | - | - |
| `ReadOps` | `readdir`, `read_bytes`, `stat` | Required filesystem behavior. |
| `NativeReadOps` | `read_stream`, `read_range`, `exists`, `find`, `du` | Optional equivalent fast paths. |
| `WriteOps` | `write`, `append`, `pwrite`, `create`, `mkdir`, `unlink`, `rmdir`, `rm_r`, `rename`, `copy`, `dir_copy`, `truncate`, `set_attrs` | Individually supported mutations. |

A native range receives `(accessor, path, index, offset, size)`; `size=None`
means through EOF. Its write twin, `pwrite`, receives `(accessor, path, data,
offset)` and keeps every byte outside the window, as pwrite(2) does. A table
with `write` and no `append` or `pwrite` gets both built from `read_bytes` and
`write`, so supply them only when the backend writes a range natively. Use `FileStat.size=None` when rendered size is unknown.
`du` supplies both size and entry enumeration. All operations enforce the same
resource scope, including direct reads of known IDs. A partial listing cannot
prove an omitted resource absent.

Supplying a write callback does not enable deletion or rename. Every generic
command is registered either way: `gzip -c` and `tar -t` run as readers, a line
that needs a missing operation answers `Operation not supported` at that
operation, and mount mode still controls whether a supported mutation may run. Set `local=True` only for host-local data;
`is_mounted` can override the default availability check.

`VFSAdapter.to_command_io()` assembles the single `CommandIO` used by commands
and filesystem ops. Builtins compile `VFSAdapter` into this table; advanced integrations can also
supply `CommandIO` directly. `overrides=` suppresses a generic command you replace, and
`commands=[...]` adds bespoke `@command` verbs.

VFS/FUSE ops are derived from the same table automatically (`make_generic_ops` under the hood): read/readdir/stat, a `glob` walk over `readdir`, plus whatever mutations the table carries. Pass `ops=[...]` only for irregular handlers (they shadow same-named derived ops), or `auto_ops=False` to opt out. A verb the op table does not carry is not served: the mount answers `Operation not supported` for it, so a backend can be partial.

The driver never sees the mount it runs under. The index store, the registered tables and the `vfs:` reference it was built from all live on the mount, built when the driver is placed and shared with any alias of the same instance.

### How the driver connects to a workspace

`BaseVFS` is the backend authoring API. Mount it once with
`Workspace({"/jira": JiraVFS(config)})`; consumers use workspace paths.

| Consumer | Connection |
| - | - |
| Shell and coreutils | `commands()` derives generic commands from the adapter; the shell supplies expansion, pipes, redirects and session policy. |
| Filesystem calls | `ops()` supplies the dispatcher's operations used by `ws.vfs`. |
| CLIs | The host registers a `CLISpec` with `register_cli`. File-oriented handlers use `inv.doors.dispatch` and `inv.doors.ns`; mounting a driver does not install an account CLI. |
| Runtimes | Runtimes wired to `RuntimeVFS` reach the same dispatcher. A process or remote runtime needs its configured mount bridge. |
| Namespace | The workspace owns symlinks, nested mounts and metadata overlays; the backend implements its own tree. |

Use `path.vfs_path` to address your backend and `path.virtual` for workspace
paths. Use the index passed to callbacks: the mount scopes it for ownership
and freshness. Put shared behavior in the adapter callbacks; an explicit
`ops=` override changes that operation, while a `commands=` override changes
that shell command. Optional mutations remain optional, and unsupported
operations report `Operation not supported`.

### Native search and backend-specific core functions

Builtin backends use the same `VFSAdapter` groups. Operation protocols live in
`mirage.vfs.types`; `CommandIO` inherits those contracts and adds command context.
Core functions that already accept the contract can be wired directly. Otherwise,
write a small async wrapper that resolves `PathSpec`, calls the resource API, and
normalizes the result. Inheriting a class does not adapt an incompatible signature.

Search is a separate optional capability:

```python theme={null}
from mirage import DuOps, SearchOps, SearchQuery

async def search(accessor, path, query: SearchQuery, index=None) -> list[str] | None:
    # Interpret query.query and validate the options your resource supports.
    ...

adapter = VFSAdapter(
    read=ReadOps(readdir=readdir, read_bytes=read_bytes, stat=stat),
    search=SearchOps(search=search),
)
```

`SearchQuery(query="recent deployments", options={"limit": 20})` carries plain
search text and backend-specific arguments. The options can hold filters, limits,
or other JSON values. `SearchOps.meta` is optional static capability metadata.
Neither regex support nor grep compatibility is required. Omitting `search`
entirely also works: MIRAGE implements grep and rg by reading files.

A resource search is not automatically used by grep or rg. To opt into that
integration, declare `SearchOps(search=search, meta={"grep": {"mode": "literal"}})`
or use `"mode": "regex"` when the backend honors regex semantics. Only this
integration interprets the `grep` namespace. Each request puts case, fixed-string,
whole-word, and basic-regex flags in `query.options["grep"]`, using snake\_case
keys in both languages. Other metadata and options belong to your resource.

A search callback returns text records: `[]` means no matches and `None` declines
the request. Errors propagate. Under the grep integration, records must be
complete rendered output lines, including path prefixes. A new filesystem search
accelerator must match scanning the rendered files. Existing backends retain
their declared output semantics, including Langfuse's summary search. Never
return a truncated result as a complete answer.

MIRAGE scans for unsupported flags, multiple operands, declined requests, and
visibility restrictions. Set `meta={"grep": {"mode": "regex", "stream": True}}`
to use native streaming during fallback scans. The hierarchy kit's
`make_search_op(detect_scope, SEARCHERS, stat)` can adapt scope-specific callbacks;
PostgreSQL, MongoDB and Langfuse use it. Custom resources can implement `search`
directly. Expose semantic or service-specific queries through a custom command
that calls the same callback with its own options.

`examples/python/other/basevfs_views.py` wires the same literal search callback
to both commands. Its counters verify that `grep -F` and `rg -F` avoid file reads,
line numbers and regexes scan, and `-i` scans after the callback returns `None`.
The example accelerates single-page searches; directory searches also decline.

Native traversal and size enumeration remain separate capabilities.
`NativeReadOps(du=DuOps(size=size, entries=entries))` supplies both halves of the
native size contract.

To make the backend constructible by name (workspace YAML, snapshots, the daemon), register it:

```python theme={null}
from mirage import register_vfs

register_vfs("jira", JiraVFS, JiraConfig)
```

or ship it as a normal package with an entry point — discovered automatically at registry-build time:

```toml theme={null}
[project.entry-points."mirage.vfs"]
jira = "mypackage.backends:JiraVFS"
```

The entry point resolves to the VFS class; declare a `CONFIG_CLS` class attribute when the constructor takes a typed config.

Neither step is needed to mount from YAML. A `vfs` value carrying a colon names the class directly, the same way a `clis` entry's `cli` value names a spec tree, so a deployment can point at a file next to the config or at a class inside an installed package:

```yaml theme={null}
mounts:
  /jira:
    vfs: ./jira.py:JiraVFS
  /wiki:
    vfs: mypackage.backends:WikiVFS
```

A relative path resolves against the config file's directory, not the server's working directory. A registry name always wins over a reference, so a name can never be reread as code. See `examples/python/other/custom_vfs.py` for a complete runnable backend in one file.

When mounting a driver built with `build_vfs` in code, preserve its loader name on the placement: `Mount(build_vfs("jira", config), vfs_ref="jira")`. YAML does this automatically. Snapshots also preserve the mount’s effective index settings. An index URL containing credentials is redacted; pass a `Mount(..., index=...)` override with fresh credentials when loading it.

Snapshots and versions rebuild a saved mount through the same door: the registered name, or the reference the config named (recorded beside the class path, since a class loaded from a script file cannot be imported back). What comes back depends on what the VFS owns. Content the VFS holds itself (an in-memory store) is mirage-owned state: override `get_state` and `load_state` to carry it, and a snapshot or a version restores the mount with that content and no override. Content that lives in a remote service is only observed: keep the default state, which says `needs_override`, set `supports_snapshot=True` and fill `FileStat.fingerprint`, and a snapshot pins what it read while `Workspace.load` asks for the live VFS back through `mounts=`. A forgotten override is a refusal to load, never a mount that comes back empty. The example shows both halves: a wiki page is written, the workspace is snapshotted, the page is changed, and the loaded workspace serves the page as it was, while a feed mount that keeps the default state is refused until the load hands it back through `mounts=`.

## Contribute a Builtin VFS

Builtins live inside the Mirage repo: one backend is four layers with one name (accessor, core, ops, VFS) plus its commands, and every layer has a TypeScript twin. Change both languages in the same PR; where they disagree, the more correct side wins.

Most of a backend is already written as a kit. Reach for one before writing a layer by hand:

| Kit | Module | For |
| - | - | - |
| API client | `mirage.core.api` (`api_request`, `cursor_items`, `TokenManager`, `RetryPolicy`) | every HTTP call: one status-to-error mapping (`error_of`), retry, pagination, OAuth refresh, and the mount's session pool |
| Hierarchy | `mirage.core.hierarchy` (`Scope`, `Slot`, `make_detect_scope`, `make_readdir`, `make_stat`, `entry_stat`, `make_read`, `DirListing`, `SearchQuery`) | an API tree of `<label>__<id>` directories: one scope table classifies every path for readdir, stat, read and search |
| Object store | `mirage.core.object_store` | a flat key space (S3-style buckets, GridFS) |
| Render | `mirage.core.render` (`json_bytes`, `jsonl_bytes`) | records rendered as the bytes the tree's files hold |

Two builtins are the references. Trello is the hierarchy kit end to end: a scope table, listers, id-addressed readers, and nested `trello <noun> <verb>` commands. Jaeger is the client shape to copy: its `_get(accessor, ...)` takes the accessor and passes `accessor.pool` itself, so no call can forget the pool. Paths are always `PathSpec` values inside the VFS; never pass a path as a raw string.

## File Structure

```text theme={null}
python/mirage/
  accessor/<name>.py         # the client handle; owns config and the session pool
  core/<name>/
    client.py                # API calls, all through mirage.core.api
    scope.py                 # SCOPES and detect_scope (hierarchy kit)
    pathing.py               # <label>__<id> names (sanitize_name, make_id_name)
    normalize.py             # the JSON a file renders
    readdir.py stat.py read.py
  ops/<name>/__init__.py     # OPS derived from the CommandIO table
  commands/builtin/<name>/
    io.py                    # VFSAdapter compiled to the shared IO table
    __init__.py              # COMMANDS
    <bespoke commands>.py    # push-downs and verbs
  vfs/<name>/
    __init__.py config.py prompt.py <name>.py
```

Tests mirror `mirage/` one to one under `python/tests/`, with no `__init__.py` in the test tree. Patch a client call where the module under test imports it.

## 1. Config, Accessor, and Registry

Define a typed Pydantic config that refuses unknown keys, and keep secrets in `SecretStr` fields. `build_vfs`, YAML and snapshots all construct through it, so a misspelled knob fails by name instead of being dropped:

```python theme={null}
from pydantic import BaseModel, ConfigDict, SecretStr


class MyConfig(BaseModel):
    model_config = ConfigDict(extra="forbid")

    token: SecretStr
    board_ids: list[str] | None = None
```

Create an `Accessor` subclass for an SDK client. For the built-in HTTP kit, subclass `SessionAccessor` from `mirage.accessor.base` and call `super().__init__()`; it owns the `SessionPool` available as `pool`. Plain `Accessor` has only a no-op `close`, so it does not allocate a pool. `BaseVFS.close()` closes the accessor (the index store is the mount's, and the mount closes it); override the accessor's `close()` for an owned client, as above. If the VFS owns additional handles, release those in its `close()` and call `super().close()`. Shared clients need an explicit ownership decision: close them in the embedding program if several mounts use them.

Add the VFS name to `VFSName`, export the config and VFS from `vfs/<name>/__init__.py`, and add a lazy `VFSEntry` to `mirage/vfs/registry.py`. The registry entry lets YAML, snapshots, and the daemon construct a builtin by name.

Keep every import at module scope. If that creates a cycle, change the dependency direction instead of adding a function-local import.

## 2. Core VFS Operations

Implement only the operations the backend supports. A read-only API-backed mount usually starts with:

* `readdir(accessor, path, index)` returning child paths.
* `read_bytes(accessor, path, index)` returning bytes.
* `stat(accessor, path, index)` returning `FileStat`.

Every client call passes `session=accessor.pool` (or takes the accessor, as Jaeger's does); without it `api_request` opens a fresh session per request. `tests/commands/test_client_calls_pass_session.py` fails a call that omits it.

A hierarchy backend writes its tree once, as `SCOPES` in `core/<name>/scope.py`, and builds the three operations from it: `make_readdir(detect_scope, listers=...)`, `make_stat(detect_scope, readdir, entry_stats=...)` and `make_read(detect_scope, readers, stat=stat)`. The kit holds these rules, and a new backend keeps them:

* A reader that reaches the API by the ids in the path slots passes `stat=` so the kit proves the file's parent through the listing first. Without it, a path outside the configured scope (Trello's `workspace_id` and `board_ids`) reads while `ls` and `stat` say it does not exist.
* A listing that is a filtered or truncated view (one page of a bounded query, a time window, a glob-scoped span) returns `DirListing(entries, partial=True)`. Cached as the whole directory, it would prove every entry it left out absent.
* `entry_stat("<id_key>", ...)` names the id under the key its path slot declares.
* `FileStat.size` is the rendered byte length or `None`, never a storage-side number. Set `sizes_always_known` only when every listed size is computed from the same payload a read renders.
* Id-addressed commands honor the same scope knobs the listing does (Trello's `commands/builtin/trello/_scope.py`).

Glob resolution is not a per-backend file: `make_generic_ops` derives a `glob` op from the table's `readdir`, capped by its `max_glob_matches`, and the mount expands patterns through it.

Use explicit types:

```python theme={null}
from mirage.accessor.base import Accessor
from mirage.cache.index import IndexCacheStore
from mirage.types import FileStat, PathSpec


async def stat(
    accessor: Accessor,
    path: PathSpec,
    index: IndexCacheStore | None,
) -> FileStat:
    ...
```

Add write, append, create, unlink, rename, or directory operations only when the backend has matching semantics. I/O stays async-native.

## 3. Ops Layer

Build `IO = VFSAdapter(read=ReadOps(...), native=NativeReadOps(...), writes=WriteOps(...)).to_command_io()` in `commands/builtin/<name>/io.py`, omitting groups the backend does not implement.

Ops are derived, not hand-written. `ops/<name>/__init__.py` builds the whole VFS/FUSE op family from the same `CommandIO` table the commands use:

```python theme={null}
from mirage.commands.builtin.my_vfs.io import IO
from mirage.ops.generic import make_generic_ops

OPS = make_generic_ops("my_vfs", IO)
```

`make_generic_ops` emits read/readdir/stat and glob plus whatever mutations the table carries: a `CommandIO` slot updates commands and ops together, and ops whose table field is `None` are omitted. Knobs mirror backend semantics, e.g. `make_generic_ops("databricks_volume", IO, mkdir_parents=True)`.

Write a dedicated op module only for an irregular handler with no generic equivalent (such as a semantic `search`), and append it to the derived list. Mark a hand-written mutation op with `write=True` so `MountMode.READ` remains a real boundary (derived ops carry this from the table).

## 4. Commands

Build the standard set with `make_generic_commands("<name>", IO, overrides={...})`; the generic command owns flag interpretation, so a backend wrapper is wiring only. `overrides` names the builders a bespoke command replaces, and a name no builder has is refused.

* Declare native text search in `VFSAdapter(search=SearchOps(...))`. The generic grep/rg builders consume it only when `search.meta["grep"]` opts in; semantic search does not opt in. The hierarchy kit supplies `make_search_op` for scope-based core functions; unsupported requests fall back to scanning. Existing wrappers may call `run_search(IO, name, ...)` when preserving a custom registration.
* A verb (`trello card create`) declares its own `CommandSpec` with every id as a flag, reads flags through `FlagView(opts.flags, spec=SPEC)` (never `flags.get(...)`), registers `write=True` for a mutation, and calls `require_mount_writable()` before the client. Specs declared inline are dumped to `.cache/spec/python/vfs_commands/` and compared with TypeScript's, and a prompt that teaches the verbs is pinned against their specs (`tests/vfs/trello/test_prompt.py`).
* Handlers take `(accessor, paths, texts, opts)`. Read `stdin`, `index`, `prefix`, and namespace facts from `opts`; do not add injected parameters or a second handler convention.
* A handler that reads its `accessor` directly is trusted host code. Admission judges the paths it is given, but no policy sees what it reads below them, so keep such a handler to its operands.

Export the final list as `COMMANDS` from `commands/builtin/<name>/__init__.py`.

## 5. VFS Class

Subclass `BaseVFS`, declare the facts as class attributes, and return the module tables from `ops()` and `commands()`. Core functions stay independent of the VFS class:

```python theme={null}
from mirage.accessor.my_vfs import MyAccessor
from mirage.commands.builtin.my_vfs import COMMANDS
from mirage.commands.config import RegisteredCommand
from mirage.commands.config import registered_commands
from mirage.ops.my_vfs import OPS as MY_VFS_OPS
from mirage.ops.registry import RegisteredOp
from mirage.vfs.base import BaseVFS
from mirage.vfs.my_vfs.config import MyConfig
from mirage.types import VFSName


class MyVFS(BaseVFS):
    name: str = VFSName.MY_VFS
    caches_reads: bool = True

    def __init__(self, config: MyConfig) -> None:
        super().__init__()
        self.config = config
        self.accessor = MyAccessor(config)

    def ops(self) -> list[RegisteredOp]:
        return MY_VFS_OPS

    def commands(self) -> list[RegisteredCommand]:
        return registered_commands(COMMANDS)
```

The mount registers both tables when the driver is placed; a driver never registers anything onto itself.

Set `caches_reads=True` only for stable, read-mostly content. Implement `get_state()` with credentials redacted (`self.config_state(self.config)`) and close any network clients in `close()`.

### Point lookups under `fresh`

A `read: fresh` mount re-stats a cached file through a throwaway store that
starts with none of the mount's rows (`ListingCheckStore`), so no cached
row answers the check.
A backend with a path lookup answers with one request for that path (Dropbox).
A backend that addresses items only by id may, after checking that the index
is a `ListingCheckStore`, read the mount's last row for
the path with `await index.hint(key)` and address one request by its id (Box);
it must then confirm from that answer alone that the item still sits at
exactly this path, and otherwise resolve the path as it always does. The hint
is a lead, never an answer: a stat built from its fields would let a stale row
pass a freshness check.

### Listing versions

Under `read: fresh`, a cached listing the running command did not write is listed again, unless the backend declares what to check it against. Declare `listing_version` (`mirage.types.ListingVersion`) as a class attribute, like `read_revalidatable`:

* `NONE` (the default): no version; every command re-lists.
* `MOUNT`: one version covers every listing of the mount (GitHub and the Hub use the head commit). A stat of the mount root through the gate's empty check store (`mirage.cache.index.ram.ListingCheckStore`) asks the backend and returns it as `FileStat.fingerprint`; through any other index it names none and reads neither the index nor the backend, since nothing reads a root fingerprint off a mount-view stat. The tree fill seeds the same value with every listing it writes (`seed(..., version=...)`).
* `FOLDER`: each listing carries its own folder's version (disk). A stat of the folder returns it as the fingerprint, and `readdir` reads it before listing and passes it to `set_dir(..., version=...)`, so a change during the scan leaves the stored version behind.

Take the version from a backend response, the same kind of token the stat answers, since the gate compares the two with `==`. Store `None` when there is nothing reliable to store; that listing re-lists.

`listings_pin` is an instance attribute (default `None`). Set it in `__init__` when the config pins the mount to something that cannot move: the lowercased ref when it is a full 40- or 64-hex commit sha. A stored listing whose version equals the pin is served with no request. The stored value still comes from a response, never from config, so a full-sha ref is served unchecked only when its listing was fetched at that sha. github.com refuses a branch or tag named with 40 or 64 hex characters, so a full-sha GitHub ref always names a commit (a GitHub Enterprise host is assumed to do the same). The Hugging Face Hub repos set no pin: a Hub revision named like a sha is checked every command.

`tests/vfs/test_listing_version.py` holds every declarer to this: the fill stores a version on the root and on a nested folder, a stat through an empty index answers the stored value, a second command sends only the expected checks and no refill, and an outside change moves the version. A new declarer adds a harness there and adds its name to the pinned roster, or the roster tests fail.

## 6. Snapshot Support

Leave `supports_snapshot=False` unless the complete drift contract is implemented:

1. `stat()` returns a stable `FileStat.fingerprint`.
2. Every read record includes the fingerprint that produced those bytes.
3. If the backend supports immutable revisions, reads consult `revision_for(path.virtual)` and record the resolved revision.

Setting the flag without recording fingerprints does not provide drift detection.

## 7. Verification

Exercise a custom backend at a nested prefix, with globs, unknown file sizes, a read-only mount, and a deliberately omitted mutation. Assert stdout, stderr, and exit status together. Use a small bounded page (three entries with five available) and request counters to check both cold and warm traversal costs. A partial `DirListing` caches positive membership until expiry; it never proves an omitted entry absent, and a subsequent directory read fetches a new page. Custom index stores may override `set_partial_dir`; the default conservatively refreshes on lookup.

* Unit tests for config validation (an unknown key refused), path layout, every VFS op, command behavior, read-only enforcement, scope enforcement, state redaction, and cleanup.
* Regenerate `spec/` with `scripts/gen_specs.py` and `typescript/scripts/gen-specs.ts`, then run `scripts/check_spec_parity.py` and `scripts/check_layout_parity.py --strict`.
* An integration target in `integ/targets.json` with cases under `integ/vfs/<name>/`, run by both hosts' runners against the same goldens; a SaaS backend gets a fake under `integ/server/`. Any change in observable shell behavior adds a case.

## Updating Existing Backends

### From `GenericVFS` and the per-driver methods

`BaseVFS` is now the one driver contract, and a driver serves only through
the tables `ops()` and `commands()` return. Nothing keeps the old spelling
alive, so a custom backend written against it changes in these places:

| Before | Now |
| - | - |
| `GenericVFS(name=..., accessor=..., io=...)` | `BaseVFS(name=..., accessor=..., io=...)`; `io` takes a `VFSAdapter` or a `CommandIO`, and a driver built from one needs a name |
| `PROMPT`, `WRITE_PROMPT`, `SUPPORTS_SNAPSHOT`, `SIZES_ALWAYS_KNOWN`, `READ_REVALIDATABLE` | `prompt`, `write_prompt`, `supports_snapshot`, `sizes_always_known`, `read_revalidatable` |
| `self.register(fn)` and `self.register_op(op)` in `__init__`, `ops_list()` | return the tables from `commands()` and `ops()`; the mount registers them when it places the driver |
| `resolve_glob(paths, prefix)` | the `glob` op `make_generic_ops` derives from `readdir`; the mount stamps each spec's `vfs_path` before calling it |
| `vfs.read_bytes(path)`, `vfs.stat(path)` and the other per-driver methods | `DriverOps(vfs)` outside a workspace, `ws.dispatch(...)` inside one |
| `vfs.index`, `set_index(...)`, `index=` on the constructor | `Workspace(index=...)` or `Mount(index=...)`; the store is the mount's (`ws.mount(prefix).index_store`) |
| `storage_id()` | `storage_location()`, which may return `None` |
| `statfs()` | `capacity()` |
| `vfs.vfs_ref` | `Mount(vfs_ref=...)`, recorded on the mount |

Mount configs now reject unknown fields. Remove PostgreSQL's `default_search_limit` and MongoDB's `default_doc_limit` and `default_search_limit`; the remaining read ceilings are `max_read_rows` and `max_doc_limit`. PostgreSQL `head`/`tail` and MongoDB `tail` report a clipped window through stderr and a nonzero status; consumers must check status before treating captured output as complete. PostgreSQL refuses whole reads over its thresholds. MongoDB streams and search no longer apply a silent result cap.

The hierarchy name helpers sanitize empty and dot-leading labels to reachable names. Regenerate stored paths from the listing instead of retaining an old sanitized label; the provider id remains the stable identifier. Keep read and mutation scope checks on every entry point, including direct ids, rather than relying on what a previous listing exposed.

PostgreSQL whole-file reads check a bounded result’s database JSON byte size before transferring rows, then check the rendered JSONL size. Database formatting can make the first check conservative; use an explicit row window when that guard refuses a read.

### Check the adapter contract

The read-contract helper accepts either an adapter or its compiled I/O table.
Supply a small known file, its parent directory, an absent sibling, and the
expected bytes. It checks listing, metadata, reads, streams, native ranges,
existence, and missing-path errors without mutating the resource.

```python theme={null}
from mirage import ReadFixture, check_read_contract

await check_read_contract(adapter, accessor, ReadFixture(
    file=file_path,
    directory=parent_path,
    missing=missing_path,
    content=b"known fixture bytes",
))
```

`check_driver_contract(vfs, fixture)` runs the same read checks against a whole
driver, through the op table `vfs.ops()` serves, the one channel a mount
dispatches to. It checks a builtin-shaped subclass as readily as a driver built
from an adapter, and always probes the `read` op's byte window. `DriverOps(vfs)`
is the table it drives: it calls a driver's ops the way a mount does, with the
accessor bound and one index store per instance, which is also how to script a
driver outside a workspace.

```python theme={null}
from mirage import DriverOps, check_driver_contract

await DriverOps(vfs).write(file_path, b"known fixture bytes")
await check_driver_contract(vfs, fixture)
```

Use disposable fixtures to test supported writes and verify read-only mounts
refuse mutations. Resource-specific tests should cover pagination, scope,
authorization failures, and query options.

Builtin VFS classes and a `BaseVFS` built from an adapter serve through the
same two tables: commands, filesystem ops, and glob expansion all come from the
backend's I/O table, and a driver carries no direct verb methods beside them.
Backend classes keep configuration, client lifecycle, storage location, watch
hooks, and snapshot behavior.

### Batch resource search

`SearchOps(search=search_one, search_many=search_many)` optionally accepts a
batch callback with `(accessor, paths, query, index)`. Use it when ranking and
`top_k` must apply once across several paths. The single-path callback remains
required. `mirage.vfs.search.search_resources` uses the batch callback when
provided and otherwise concatenates single-path results; a declined request
raises an error rather than reporting no matches.

Chroma, Dify, Qdrant, LanceDB, and Mem0 use this capability for their existing
resource search commands. Their options remain backend-specific: for example,
`top_k`, `threshold`, and `method`. These declarations do not opt into grep.

### Start from a packaged example

The `mirage-vfs-authoring` skill in the Mirage plugin includes self-contained
Python and TypeScript starters and `scripts/new_adapter.py`. It creates an
adapter file in your project, refuses to overwrite existing files, and includes
the contract check plus a mounted shell smoke test. Replace its fixture client
with your resource API, then expand capabilities as needed.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.