> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mirage.strukto.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# MongoDB

> Mount MongoDB databases, collections, and documents as a Mirage filesystem with JSON access for Python agents.

The MongoDB VFS exposes MongoDB databases, collections, and documents
as a virtual filesystem mounted at some prefix such as `/mongodb/`.

For connection setup, see [MongoDB Setup](/home/setup/mongodb).

## Config

```python theme={null}
import os

from mirage import MountMode, Workspace
from mirage.vfs.mongodb import MongoDBConfig, MongoDBVFS

config = MongoDBConfig(
    uri=os.environ["MONGODB_URI"],
    max_doc_limit=5000,
    elide_fields={"sample_mflix.movies": ["plot_embedding"]},
)
vfs = MongoDBVFS(config=config)
ws = Workspace({"/mongodb": vfs}, mode=MountMode.READ)
```

| Config field | Required | Default | Description |
| - | - | - | - |
| `uri` | yes | | MongoDB connection URI |
| `databases` | no | | List of database names to mount (omit for all) |
| `max_doc_limit` | no | 5000 | Most documents one `tail -n K` returns; past it, a notice and exit 1 |
| `elide_fields` | no | `{}` | `{"<db>.<coll>": ["field", "nested.path"]}`; listed fields are dropped from `documents.jsonl` |

## Filesystem Layout

The mount mirrors MongoDB's `cluster → database → collection / view`
model. The database directory always appears in the path, even when
`databases` filters to a single entry.

```text theme={null}
/mongodb/
  <database>/
    database.json                  # collection and view names
    collections/
      <collection>/
        schema.json                # sampled types + indexes + validator
        documents.jsonl            # streamed BSON Extended JSON
    views/
      <view>/
        schema.json                # view-aware (no indexes / validator)
        documents.jsonl            # streamed, same format as collections
```

Example:

```text theme={null}
/mongodb/
  sample_mflix/
    database.json
    collections/
      movies/
        schema.json
        documents.jsonl
      comments/
        schema.json
        documents.jsonl
    views/
      top_rated_movies/
        schema.json
        documents.jsonl
```

Discovery lists databases, collections, views, and their fixed file names.
It does not enumerate one filesystem entry per document. Schema inspection
is separate from bulk document reads, so discovering a collection does not
require exporting it. These paths and read contracts are shared by Python
and TypeScript.

### documents.jsonl

One JSON object per line, encoded with BSON
[Relaxed Extended JSON](https://www.mongodb.com/docs/manual/reference/mongodb-extended-json/).
BSON-specific types round-trip through their canonical `$` wrappers:

```json theme={null}
{"_id":{"$oid":"573a1390f29313caabcd4135"},"title":"Casablanca","released":{"$date":"1943-01-23T00:00:00Z"},"runtime":102}
```

`cat` / `head` / `tail` / `grep` / `jq` can read this file through the same
streaming interface. The cursor fetches documents in batches; `head` can stop
after a prefix. Commands that need the whole input, such as `sort` or `jq -s`,
can still buffer it. Use bounded reads when exploring a large collection.

Because the file is rendered on demand, `stat` / `ls -l` report no size and
`du` returns `0` for it (computing the real size would require rendering the
whole collection). Use `wc -c documents.jsonl` for the actual rendered byte
count, or `wc -l documents.jsonl` for an exact document count (which can scan
the collection). `stat -c %s` prints `-` for an unknown size and `0` only for
a known empty file. Missing dates also display as `-`.

Long directory listings include a `total` row. GNU `ls` reports allocated disk
blocks there; the VFS does not expose allocation metadata, so Mirage prints
`total ?` for nonempty listings and `total 0` for empty listings. This is an
intentional divergence from GNU output, not a sum of rendered file sizes.

`stat` checks existence and type without counting documents, sampling fields,
or fetching indexes. This keeps `ls`, `find`, and opening a stream independent
of the collection size. The same unknown-size convention applies to Postgres
row exports and other rendered VFS files.

### schema.json

Generated from the first 100 documents in `_id` order. Includes:

* field path → observed BSON type frequencies (nested paths unioned across the sample)
* index definitions from `listIndexes`, without volatile access counters
* the `$jsonSchema` validator if one is registered

Repeated reads of unchanged data produce identical bytes. Schema reads do
not include document counts or `$indexStats`, so reading a schema cannot
change its own content. Views skip indexes.

### database.json

Lists collection and view names using two catalog listings, without a
per-collection count. Useful for `cat /mongodb/<db>/database.json` to get an overview
without recursing into each entity.

### Elided fields

Fields listed under `elide_fields` are dropped entirely from
`documents.jsonl` output. The type stays documented in `schema.json`, so
heavy fields (embeddings, large binary, raw text blobs) can be hidden
from agent reads without losing the schema signal:

```python theme={null}
config = MongoDBConfig(
    uri=os.environ["MONGODB_URI"],
    elide_fields={
        "sample_mflix.movies": ["plot_embedding"],
        "rag.docs": ["metadata.embedding"],
    },
)
```

Nested paths use dot notation. Elision applies to both `cat` (one-shot
streaming reads) and `tail -f` (live change-stream follows).

## Streaming and Limits

`cat`, `grep`, `head`, and `tail -f` consume documents lazily through a
batched PyMongo async cursor (or change stream); the consumer cancels to stop
fetching. There is no truncation notice because nothing is forced into
memory ahead of the consumer.

| Command | Behavior |
| - | - |
| `cat` | Streams the whole collection sorted by `_id`; pipe to `head` to cap |
| `head -n K` | Streams from `_id` order and stops after K documents |
| `tail -n K` | Server-side sort + limit; a K past `max_doc_limit` stops there, notes it on stderr, exits 1 |
| `tail -f` | Opens a change stream; yields each new insert as a JSONL line |
| `grep` (file level) | Streams from `documents.jsonl`; supports `-m` for short-circuit |
| `grep` (collection/db level) | Every file under the directory, each read as `cat` renders it; no result cap |
| `jq` | Inherits the streaming `cat` source |
| `wc` | Uses `countDocuments()` server-side; zero download |
| `stat` | Existence and type only; no counts or index queries |

## Smart Commands

### grep at different scopes

`grep` uses MongoDB's query engine at directory scopes instead of
streaming all documents through the regex pipeline:

```bash theme={null}
# FILE level - streams documents.jsonl, runs the regex locally
grep "Godfather" "/mongodb/sample_mflix/collections/movies/documents.jsonl"

# COLLECTION level - server-side query against the collection
grep "Godfather" "/mongodb/sample_mflix/collections/movies/"

# DATABASE level - searches across every collection in sample_mflix
grep "Godfather" "/mongodb/sample_mflix/"

# ROOT level - fans out across every mounted database
grep "Godfather" "/mongodb/"
```

At collection or higher scope the VFS picks the best server-side
strategy from the indexes available:

1. **Text index exists** → uses `$text` (ranked by relevance)
2. **Atlas Search index exists** → uses `$search` (fuzzy, Lucene-based)
3. **Neither** → falls back to `$regex` on sampled string fields

Scope detection is handled by `mirage/core/mongodb/scope.py`.

### head / tail / tail -f

`tail` uses server-side `sort` + `limit`. A count past `max_doc_limit`
prints the last `max_doc_limit` documents, then says on stderr that the
output is incomplete and exits 1, rather than printing fewer lines than
asked in silence. `head` streams in `_id` order and stops after the
count. `tail -f` opens a Mongo change stream
filtered to insert events and yields each new document as a JSONL line
in the same format as `cat`:

```bash theme={null}
# First 10 docs (sorted by _id ascending)
head -n 10 "/mongodb/sample_mflix/collections/movies/documents.jsonl"

# Last 10 docs (sorted by _id descending)
tail -n 10 "/mongodb/sample_mflix/collections/movies/documents.jsonl"

# Live-follow new inserts; consumer cancels to stop
tail -f "/mongodb/sample_mflix/collections/movies/documents.jsonl"
```

`tail -f` requires the cluster to be a replica set; Atlas already
satisfies this. Views fall through to the non-streaming path because
change streams aren't defined on views.

## Cache

The MongoDB VFS uses `IndexCacheStore` (same as RAM/S3/disk/GitHub)
for listings: database names, collection names, and document counts.

Document content is **not** cached. The VFS leaves `caches_reads`
at its default of `False`, so `cat`, `grep`, `head`, and `tail` always
query the live collection
instead of serving a stored snapshot. This keeps reads consistent with a
mutable database and ensures `tail -f` follows the live change stream
rather than replaying cached bytes.

## Example

```python theme={null}
import asyncio
import os

from dotenv import load_dotenv

from mirage import MountMode, Workspace
from mirage.vfs.mongodb import MongoDBConfig, MongoDBVFS

load_dotenv(".env.development")

config = MongoDBConfig(uri=os.environ["MONGODB_URI"])
vfs = MongoDBVFS(config=config)


async def main():
    ws = Workspace({"/mongodb": vfs}, mode=MountMode.READ)

    # List all databases
    r = await ws.shell("ls /mongodb/")
    print(await r.stdout_str())

    # List the entities under a database (database.json, collections/, views/)
    r = await ws.shell("ls /mongodb/sample_mflix/")
    print(await r.stdout_str())

    # List collections
    r = await ws.shell("ls /mongodb/sample_mflix/collections/")
    print(await r.stdout_str())

    # Read first 5 movies
    r = await ws.shell(
        'head -n 5 "/mongodb/sample_mflix/collections/movies/documents.jsonl"')
    print(await r.stdout_str())

    # Read last 5 movies
    r = await ws.shell(
        'tail -n 5 "/mongodb/sample_mflix/collections/movies/documents.jsonl"')
    print(await r.stdout_str())

    # Inspect the sampled schema and indexes
    r = await ws.shell(
        'cat "/mongodb/sample_mflix/collections/movies/schema.json"')
    print(await r.stdout_str())

    # Extract titles with jq
    r = await ws.shell(
        'jq -r ".title" "/mongodb/sample_mflix/collections/movies/documents.jsonl"'
    )
    print(await r.stdout_str())

    # Search across a database (uses MongoDB query engine)
    r = await ws.shell('grep "Godfather" "/mongodb/sample_mflix/"')
    print(await r.stdout_str())

    # Count documents (server-side, no download)
    r = await ws.shell(
        'wc -l "/mongodb/sample_mflix/collections/movies/documents.jsonl"')
    print(await r.stdout_str())

    # View the database overview without recursing
    r = await ws.shell('cat "/mongodb/sample_mflix/database.json"')
    print(await r.stdout_str())


if __name__ == "__main__":
    asyncio.run(main())
```

## Runnable examples

Three working examples live under `examples/python/mongodb/`:

* [`mongodb.py`](https://github.com/StruktoAI/mirage/blob/main/examples/python/mongodb/mongodb.py) — agent-shell workflow: `ls`, `tree`, `cat`, `head`, `tail`, `wc`, `stat`, `grep`/`rg` at every scope, `jq`, `find`, `cd` + relative paths.
* [`mongodb_vfs.py`](https://github.com/StruktoAI/mirage/blob/main/examples/python/mongodb/mongodb_vfs.py) — in-process VFS: `os.listdir` and `open()` walk every readdir level (root, database, `collections/`, `views/`, entity) and read `database.json`, `schema.json`, `documents.jsonl` (collection + view).
* [`mongodb_fuse.py`](https://github.com/StruktoAI/mirage/blob/main/examples/python/mongodb/mongodb_fuse.py) — same coverage as the VFS example, but the tree is mounted as a real filesystem so other processes can `cat`/`ls`/`head` the mountpoint directly.

All three default to the `mirage_test` database seeded by `python/scripts/seed_mongodb_test.py`.

## Finding IDs

`_id` is serialized as `{"$oid": "..."}` under Extended JSON:

```bash theme={null}
# List the first 10 ObjectId values
jq -r '._id["$oid"]' \
  "/mongodb/sample_mflix/collections/movies/documents.jsonl" | head -n 10

# Find a specific document by ID
grep "573a1390f29313caabcd42e8" \
  "/mongodb/sample_mflix/collections/movies/documents.jsonl"

# Extract a few fields together
jq -r '"\(._id["$oid"]) \(.title) \(.year)"' \
  "/mongodb/sample_mflix/collections/movies/documents.jsonl"
```

## Working with Large Collections

Streaming is the default; reach for these patterns when you want to keep
the round trip small:

```bash theme={null}
# Document count (server-side, no download)
wc -l "/mongodb/sample_mflix/collections/comments/documents.jsonl"

# Most recent documents (sorted by _id desc)
tail -n 10 "/mongodb/sample_mflix/collections/comments/documents.jsonl"

# Server-side query at collection scope avoids streaming the whole jsonl
grep "love" "/mongodb/sample_mflix/collections/comments/"

# Drop heavy fields via elide_fields, then take the slice you want
head -n 20 "/mongodb/sample_mflix/collections/movies/documents.jsonl"
```

Hide embeddings or large blobs from agent reads with `elide_fields`;
`schema.json` still documents the original type so the agent can decide
when to ask for the raw bytes through a different path.

## Shell Commands

Standard commands available on the mounted MongoDB tree:

| Command | Notes |
| - | - |
| `ls` | List databases, collections, views, and per-entity files |
| `cat` | Stream `documents.jsonl` (or read `schema.json` / `database.json`) |
| `head` / `tail` | `head` streams and stops early; `tail` sorts and limits server-side, stopping loudly at `max_doc_limit` |
| `tail -f` | Live-follow new inserts via Mongo change stream |
| `grep` / `rg` | A directory operand searches every file under it (documents, schema.json, database.json, views); a file streams |
| `jq` | Query JSON; runs once per line of a JSONL file (`-s` for all) |
| `wc` | Smart: uses `countDocuments()` server-side |
| `stat` | Existence and type only; no counts or index queries |
| `find` | List databases/collections/views with `-name`, `-maxdepth` |
| `tree` | Directory tree view |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.