What It Does
EveryWorkspace ships with a two-layer cache so repeated work against remote backends (S3, GDrive, Slack, …) hits local state instead of the network:
- Index cache. Listings and metadata. The first directory walk hits the API; subsequent ones serve from the index until the TTL expires.
- File cache. Object bytes. The first read streams from origin; later pipelines read from cache.
read: in its mount block. Under read: fresh, cached file bytes are checked against a backend freshness probe that no cached index row answers (a backend with no path lookup, such as Box, may use the mount’s row only for the id it addresses its one request to). Within one shell command, the first probe of a path supplies its answer to later cache checks and compatible command stats. A write on that mount, the cache clear that follows an external program or a remote runtime line, or a re-list that finds the path gone retires the answer, including a probe still in flight. Another command obtains its own answer, so an outside change after the first probe can remain unseen for the rest of that command. Reads outside a command do not reuse probe answers. A changed fingerprint triggers a fresh read; a deleted file is evicted. Under read: bounded — the default — cached bytes are trusted without a check, for as long as the mount’s ttl: allows. Each listing write is also capped by the mount’s ttl, under either policy. Existing Redis listings retain their stored expiry until rewritten.
fresh is refused at mount time on a backend that cannot honour it: one that does not cache reads has no gate to check at, and one whose stat and read stamp different kinds of token could never match. A mount that asked to revalidate and quietly did not is the failure the refusal exists to prevent, so the policy is never silently downgraded. See the YAML reference for the keys and the VFS matrix for which backends accept fresh today.
The check happens wherever cached bytes are served, not only for the files a command names: a recursive walk, a glob and a cross-mount fan-out each revalidate every file whose content they read. Each check compares the cached fingerprint, reusing only an eligible answer from its own command; otherwise it asks the backend. One probe can require several requests on a backend that resolves a path through parent listings. One exemption, deliberate: the size a render-dependent backend cannot report for itself is filled from the cache without a check, so a metadata walk (ls -lR, du) can report a size from a stale render. Content is verified against that probe wherever it is served; that one derived number is not.
When the check cannot be made — the backend supplies no fingerprint, registers no stat at all, or the probe itself fails — the cached copy is dropped and the read fetches current bytes. Bytes that could not be verified are never served, but a transient probe failure does not fail the command either: the cold read that follows reports any real backend failure in the command’s own voice, and one flaky stat during a recursive walk costs that file a refetch rather than ending the walk.
A fresh directory listing proves a child absent only when that child is not listed. If a listed child’s metadata has been evicted independently, object-store stat fetches it from the backend.
A read that resolves a filetype renderer is never served from the file cache, and its rendered bytes are never kept there. A filetype renderer is a read op scoped to an extension, whether the VFS ships it (Google Docs, Sheets and Slides do) or it is registered beside the VFS. The file cache holds what the VFS’s own IO returns, whether a command read it or a cold read through the dispatcher kept it. For a renderer registered beside the VFS, those bytes need not be what the renderer returns, and the dispatcher does not tell the two kinds apart, so a shipped renderer whose bytes match the cache refetches too.
The rule covers any read through the dispatcher: ws.vfs.read, FUSE, a cross-mount cp or relay and < file always run the renderer. A cross-mount relay such as cp, sed or diff does not keep the rendered bytes it read either. Under read: fresh the cached entry is still checked first, so a path the backend reports gone fails as it does for any other read. The cost is that a renderer read is never warm: each one runs the renderer, and each range of a ranged read (FUSE reads in chunks) runs it again.
It does not cover:
- commands that read a file operand through the VFS’s own IO (
cat,grep,head,wc), a symlinked operand included, which keep their cache hits; - a user override of the plain
readop (one registered with no filetype), since only filetype-scopedreadops are checked; - a user-registered command that reports rendered bytes in its result’s
cachelist: the file cache stores whatever a command reports, and command reads such ascatthen serve it.
Listings under fresh
fresh also checks cached listings. A live cached listing (one still inside the mount’s ttl) is served when one of these holds, checked in this order; otherwise the folder is listed again:
- The running command wrote it. A read outside any command (FUSE,
ws.vfs, an agent’s file tools) trusts a listing written in the last second instead. - The mount is a GitHub mount pinned to a full 40- or 64-hex commit sha, and the listing was fetched at that sha. This costs no request: github.com refuses a branch or tag named with 40 or 64 hex characters, so such a ref always names a commit (a GitHub Enterprise host is assumed to do the same). A Hugging Face Hub mount never pins: a full-sha revision is checked every command like a branch, since a branch or tag named like the sha could take the name, and mirage does not assume which one the Hub resolves.
- The backend versions its listings, and the version stored with the listing still matches the backend’s.
- GitHub, HF Models, HF Datasets, HF Spaces: one version for the whole mount, the head commit the ref or revision resolves to. On the Hub a mount with a
key_prefixstores that commit joined with its key prefix, so mounts of different subtrees of one repository that share an index never share a version. A command sends one small check for the mount (GitHub: a shallowgit/trees/{ref}; the Hub:revision/{rev}?expand[]=sha, about 110 bytes) and then serves every cached listing of the mount at that commit. A moved head re-lists, which on these backends means one full tree fetch. A GitHub repo whose recursive tree comes back truncated stores no version and still re-lists each folder per command. - Disk: each folder’s own version, from its device, inode, change time and modified time. A command pays one local stat per listed folder instead of a scan. A folder changed in the last 2 seconds stores no version and re-lists until it has been quiet that long. This assumes a local POSIX filesystem; set
folder_versions: falsefor a network or FUSE root, where a folder’s change time may not move (see Disk). - Every other backend: no version. A cached listing is re-listed once per command, then trusted for the rest of that command. Google Drive and OneDrive, for example, expose no folder value that moves when a folder’s entries change.
bounded mounts never check.
Stores
Each layer is a pluggable store with two built-ins:- RAM (default): in-process, zero setup, 512 MB file cache. Best for single-process apps and notebooks.
- Redis: shared across workers, processes, and machines. Best for serverless, multi-replica services, or for cache state that survives restarts.
CacheConfig(limti="1MB") raises rather than building the defaults, and so does a Redis-only field such as url on a RAM config. Python refuses when the config is built, TypeScript when the Workspace is built. TypeScript also takes a field’s snake_case spelling (key_prefix for keyPrefix), as its VFS configs do, but not both at once.
In YAML, an index: block’s values are checked when the file is loaded, at the workspace level and in a mount block alike: url and key_prefix are strings and ttl is a number of seconds, fractions included, so a quoted ttl (including one filled in from ${VAR}) or a boolean is refused. In code, Python’s IndexConfig still converts a quoted or boolean ttl.
Index TTL
How long a cached listing lives comes from the first of these that applies:- A per-mount index: the mount block’s
index:in YAML (a mount’s own index),Mount(vfs, index=IndexConfig(ttl=60))/ws.add_mount(prefix, vfs, index=...)in Python, ornew Mount(vfs, { index: { ttl: 60 } })/ws.addMount(prefix, vfs, mode, read, vfsRef, index)in TypeScript. It replaces the workspace index whole, so a RAM mount index under a Redis workspace index keeps that mount’s listings in RAM. A second mount of the same VFS instance shares the first mount’s index, through every door, so its ownindexis not used. - The workspace index, the
index:block in YAML or theindexoption in code. Itsttlapplies to every mount without its own index, and is 600 seconds when left out, including on backends whose own TTL is shorter or zero. - The backend’s own
index_ttl(indexTtlin TypeScript), on a RAM index, when no index is configured at all:
Mounts whose Redis index names the same
url and key_prefix share one key space, keyed by virtual path, whether the index is the workspace’s or a mount’s own. Within one workspace the mount prefixes keep them apart; give each workspace its own key_prefix.
Whichever applies, each listing write is capped by the mount’s ttl:, which is 600 seconds unless the mount sets one. Listings already in a shared Redis index keep their stored expiry until rewritten or invalidated, so a workspace or process that starts with a lower ttl: can still serve them; an unmount or remount clears the mount’s listings. A 24-hour backend TTL therefore takes effect only on a mount whose ttl: is raised to match; in YAML a mount’s ttl: needs a read: beside it.
Eviction & Limits
The two layers are bounded differently:
Raising the file limit keeps more bytes warm at the cost of memory; lengthening the index TTL reduces API walks at the cost of staleness, up to the mount’s
ttl.
Miss/Hit Lifecycle
Metadata Index Contract
The RAM and Redis metadata indexes share the same lookup states: a directory never listed isNOT_FOUND, a cached empty directory is a successful empty
listing, and a listing at or past its deadline is EXPIRED. Entry metadata
remains readable after a directory expires. invalidate() expires listings
without deleting their metadata; invalidate_dir / invalidateDir removes a
listing and its direct children’s entries, and prefix invalidation removes a
whole subtree. clear() removes all index data.
Direct backend lookups trust cached metadata only when a fresh parent listing
includes the path. Partial or filtered listings may store individual entries
without a complete parent; these entries require another parent refresh on
each direct lookup. The old target is removed before refreshing so an omitted
or renamed item cannot survive as an orphaned cache hit.
Concurrent lookups sharing an index serialize each parent’s refresh through
the final entry lookup. GitHub and Hugging Face snapshot readers hold the same
guard for the whole mount, so one reader cannot remove a snapshot another is
still reading. These guards coordinate tasks within one process; unrelated
parent directories and mounts can refresh concurrently.
seed() merges snapshots by path and copies their child lists. Redis queues
these synchronous calls and flushes them before the next index operation or
close(); clear() discards queued snapshots. Multiple seeds accumulate, and
failed flushes remain queued for retry. entries() returns all stored metadata,
including entries whose directory listing has expired.
Every Redis index key starts with the config’s key_prefix, mirage:index:
by default, so the default layout is
mirage:index:mirage:idx:entry:<absolute-path>. Redis stores each entry as
the IndexEntry JSON under {key_prefix}mirage:idx:entry: and each
directory’s children and deadline as the IndexDirectory JSON under
{key_prefix}mirage:idx:directory:. Both are the documents pydantic writes, snake_case
and every field, and TypeScript writes and reads the same bytes
(IndexEntry.toJSON / IndexEntry.fromJSON), so one Redis serves both
languages and a row that does not parse is an error rather than a miss.
Invalidation uses permanent keys beside the payload:
{key_prefix}mirage:idx:generation for the whole store and
{key_prefix}mirage:idx:generation:<absolute-directory> for each directory.
Each listing records both tokens, and one MGET reads the listing and its two
current tokens. A missing or changed token makes the listing EXPIRED.
New tokens are unique so eviction or a refill cannot revive an old listing.
Competing initializers retain their attempted tokens; a worker that loses
initialization incurs a cache refill instead of adopting a later generation.
Global invalidation rotates the global token atomically. Directory invalidation
removes that directory’s token; subtree invalidation removes tokens under that
path, preserving fresh listings outside a scoped invalidation.
Like RAM, these records retain expired listings until explicit removal; they do
not use Redis TTL deletion. Redis server eviction can still turn any cached
record into a miss. A listing is served as stored even when one of its
children’s rows was evicted; GitHub and the Hub refill the tree once when a
stat or read of that listed child finds no row, and a name the listing does
not hold stays absent with no refill. Configure Redis maxmemory for
server-side size limits.
The layout is not versioned and earlier layouts are not read. Upgrade every
worker that shares an index together, and flush the key prefix first
(clear(), or delete the keys under it), so no worker opens a row another
release wrote.
The Redis file cache likewise uses server-side limits: nothing is evicted
client-side (its limit caps each file’s background fill when
max_drain_bytes is unset, and bounds max_drain_bytes), so cap
memory with maxmemory and an LRU maxmemory-policy. File bytes and metadata occupy separate
keys, so server eviction can remove either independently: a freshness check
alone does not guarantee the bytes remain cached.
Relationship To Snapshots
The file cache is exactly what a snapshot serializes:ws.snapshot() writes the cached bytes for every touched path into the tar, and Workspace.load() restores them into the file cache so a replayed run reads from local state. The index cache is not snapshotted; it rebuilds lazily after load.