Skip to main content
The HF Datasets VFS mounts a Hugging Face Dataset repo at some prefix such as /ds/. All reads are lazy: only the bytes you actually cat/head get transferred. For credential setup, see HF Datasets Setup.

Install

Config

HfDatasetsConfig takes repo_id in namespace/dataset-name form plus an optional access token. Public datasets need no token.

Reading, not writing

This mount is read-only, the way a github mount is. A Hub write is a commit, and a POSIX write cannot say where a commit ends, so echo >, rm, cp and mv are refused here rather than silently making one commit per file. The hf CLI is the write half: hf download --local-dir puts a copy on a ram or disk mount, which is an ordinary writable filesystem, and hf upload sends it back as a single commit.

Listings under fresh

Under read: fresh, a cached listing is checked against the commit the revision resolves to: one small check per command (revision/{rev}?expand[]=sha), after which every cached listing of the mount at that commit is served. A mount whose revision is a full 40- or 64-hex commit sha is checked the same way, with that one small request per command: a branch or tag named like the sha could take the name, and mirage does not assume which one the Hub resolves. With a key_prefix, the stored version is that commit joined with the prefix, so mounts of different subtrees sharing an index never share a version. See listings under fresh.

Filesystem Layout

Maps dataset repo files to virtual paths under the mount prefix. For example, if dataset AlienKevin/SWE-ZERO-12M-trajectories contains:
Then mounting at /ds/ exposes:

Example

Shell Commands

Every read command in HF Buckets’ set works here, as do the text processing and path utilities, which only read. What does not is the File Operations group: this mount is read-only, so rm and touch are refused, as is any command asked to write into it.

Cache

Uses IndexCacheStore with index_ttl = 86_400 (one day), capped by the mount’s ttl: (600 seconds unless raised; see index TTL). Directory listings are cached and populate file-size/type entries for stat’s fast path. The first listing fills the index for the whole mount: one revision request for the commit, then the recursive tree walked at that commit, so a readdir + per-entry stat (which ls, FUSE getattr, and most shell commands trigger) costs those requests rather than one per entry.

Use Cases

  • AI agents inspecting datasets: Mount, browse the README, read byte ranges from large shards without downloading the whole dataset
  • Dataset triage: ls, stat, find to see what’s in a repo before committing to a full local copy
  • Sandboxed access: Pin a revision for reproducibility