/ds/.
All reads are lazy: only the bytes you actually cat/head get transferred.
For credential setup, see HF Datasets Setup.
Install
Config
HfDatasetsConfig takes repo_id in namespace/dataset-name form plus an
optional access token. Public datasets need no token.
Reading, not writing
This mount is read-only, the way agithub mount is. A Hub write is
a commit, and a POSIX write cannot say where a commit ends, so echo >,
rm, cp and mv are refused here rather than silently making one commit
per file. The hf CLI is the write half: hf download --local-dir
puts a copy on a ram or disk mount, which is an ordinary writable filesystem,
and hf upload sends it back as a single commit.
Listings under fresh
Under read: fresh, a cached listing is checked against the commit the
revision resolves to: one small check per command
(revision/{rev}?expand[]=sha), after which every cached listing of the
mount at that commit is served. A mount whose revision is a full 40- or
64-hex commit sha is checked the same way, with that one small request
per command: a branch or tag named like the sha could take the name, and
mirage does not assume which one the Hub resolves. With a
key_prefix, the stored version is that commit joined with the prefix, so mounts
of different subtrees sharing an index never share a version. See
listings under fresh.
Filesystem Layout
Maps dataset repo files to virtual paths under the mount prefix. For example, if datasetAlienKevin/SWE-ZERO-12M-trajectories contains:
/ds/ exposes:
Example
Shell Commands
Every read command in HF Buckets’ set works here, as do the text processing and path utilities, which only read. What does not is the File Operations group: this mount is read-only, sorm and touch are refused, as is any command asked to write into it.
Cache
UsesIndexCacheStore with index_ttl = 86_400 (one day), capped by the
mount’s ttl: (600 seconds unless raised; see
index TTL). Directory listings are cached and populate file-size/type entries for stat’s fast
path. The first listing fills the index for the whole mount: one revision
request for the commit, then the recursive tree walked at that commit, so a
readdir + per-entry stat (which ls, FUSE getattr, and most shell
commands trigger) costs those requests rather than one per entry.
Use Cases
- AI agents inspecting datasets: Mount, browse the README, read byte ranges from large shards without downloading the whole dataset
- Dataset triage:
ls,stat,findto see what’s in a repo before committing to a full local copy - Sandboxed access: Pin a
revisionfor reproducibility