3.9 KiB
source_url, ingested, sha256
| source_url | ingested | sha256 |
|---|---|---|
| https://github.com/future-file-format/F3 | 2026-06-29 | 7e4e455b1bdb0733467b76d96e3d27f23857716766b9093c7922de9a08c23b1a |
F3: The Open-Source Data File Format for the Future
Source: https://github.com/future-file-format/F3 Homepage / paper DOI: https://dl.acm.org/doi/10.1145/3749163 Repository inspected: future-file-format/F3, commit bd92506 on main GitHub metadata observed on 2026-06-29: Rust, MIT license, 727 stars, 25 forks, 1 open issue, not archived.
Repository README summary
F3 is a data file format designed for efficiency, interoperability, and extensibility. The README frames it as a way to fix layout shortcomings of last-generation formats such as Parquet while preserving future-proof interoperability through embedded Wasm decoders.
The repository explicitly says this is a research prototype for verifying the paper's ideas and should not be used in production.
Important directories listed by the project:
format: FlatBuffer definition of the file format.fff-poc: main code for the F3 format, referencingfff-core,fff-encoding,fff-format, andfff-ude-wasm.fff-bench: benchmarks and experiments from the paper.fff-ude*: User-Defined-Encoding and Wasm decoding implementation.scriptsandexp_scripts: experiment scripts.
Implementation notes from local inspection
The Rust workspace contains the core proof of concept, encoding support, FlatBuffers format generation, User-Defined-Encoding support, Wasm support, benchmark code, and sample Wasm libraries. Rough source size from the inspected tree:
fff-poc: 39 files, 9,390 lines.fff-core: 6 files, 917 lines.fff-encoding: 10 files, 1,652 lines.fff-format: 3 files, 87 lines.fff-ude: 3 files, 651 lines.fff-ude-wasm: 7 files, 1,735 lines.fff-bench: 46 files, 7,454 lines.wasm-libs: 22 files, 1,351 lines.
format/File.fbs describes the file layout as row groups containing IOUnits/Chunks, which contain EncUnits. EncUnits are contiguous and can carry encoding metadata; an encoding may point to a Wasm binary stored in the same file. The file also has a section for Wasm binaries, row group metadata, optional sections, a footer, and a fixed-size postscript with metadata size, footer size, compression type, checksum type, data checksum, schema checksum, version, and the F3 magic value.
The format definition says optional metadata sections currently store Wasm binaries, but could also hold things like column UUIDs for schema evolution or zone maps for predicate pushdown.
fff-poc/src/options.rs shows default writer settings: 8 MiB IOUnit size, 64 Ki rows per encoding unit, xxhash checksum, one row group by default, encoder dictionary by default, no per-IOUnit checksum by default, and uncompressed EncUnits by default.
fff-poc/src/writer.rs builds writers around Arrow RecordBatch input, logical column encoders, row group metadata, shared dictionary context, checksum handling, and a Wasm writing context. Comments note that parts of the API are still research-oriented and some custom encoding constraints are not fully general.
fff-poc/src/reader/mod.rs includes a newer FileReaderV2 with projection and selection handling, metadata buffer loading from the postscript/footer, shared dictionary cache, checksum options, and Wasm reading context.
Reproduction and caveats
The documentation includes paper reproduction steps for metadata overhead, compression ratio, decompression speed, random access, dictionary scope, Wasm microbenchmarks, Wasm size, Wasm decoding time, and checksum overhead. Several experiments depend on external forks or code modifications, and one section says a one-off script is not yet provided.
The README says the project was only tested on an Intel machine with Debian 12. Build instructions include submodule initialization, a Debian setup script, cargo build -p fff-poc, and cargo test -p fff-poc.