Files
llm-wiki/raw/articles/f3-file-format-2026.md
T
2026-06-30 01:24:06 +09:00

3.9 KiB

source_url, ingested, sha256
source_url ingested sha256
https://github.com/future-file-format/F3 2026-06-29 7e4e455b1bdb0733467b76d96e3d27f23857716766b9093c7922de9a08c23b1a

F3: The Open-Source Data File Format for the Future

Source: https://github.com/future-file-format/F3 Homepage / paper DOI: https://dl.acm.org/doi/10.1145/3749163 Repository inspected: future-file-format/F3, commit bd92506 on main GitHub metadata observed on 2026-06-29: Rust, MIT license, 727 stars, 25 forks, 1 open issue, not archived.

Repository README summary

F3 is a data file format designed for efficiency, interoperability, and extensibility. The README frames it as a way to fix layout shortcomings of last-generation formats such as Parquet while preserving future-proof interoperability through embedded Wasm decoders.

The repository explicitly says this is a research prototype for verifying the paper's ideas and should not be used in production.

Important directories listed by the project:

  • format: FlatBuffer definition of the file format.
  • fff-poc: main code for the F3 format, referencing fff-core, fff-encoding, fff-format, and fff-ude-wasm.
  • fff-bench: benchmarks and experiments from the paper.
  • fff-ude*: User-Defined-Encoding and Wasm decoding implementation.
  • scripts and exp_scripts: experiment scripts.

Implementation notes from local inspection

The Rust workspace contains the core proof of concept, encoding support, FlatBuffers format generation, User-Defined-Encoding support, Wasm support, benchmark code, and sample Wasm libraries. Rough source size from the inspected tree:

  • fff-poc: 39 files, 9,390 lines.
  • fff-core: 6 files, 917 lines.
  • fff-encoding: 10 files, 1,652 lines.
  • fff-format: 3 files, 87 lines.
  • fff-ude: 3 files, 651 lines.
  • fff-ude-wasm: 7 files, 1,735 lines.
  • fff-bench: 46 files, 7,454 lines.
  • wasm-libs: 22 files, 1,351 lines.

format/File.fbs describes the file layout as row groups containing IOUnits/Chunks, which contain EncUnits. EncUnits are contiguous and can carry encoding metadata; an encoding may point to a Wasm binary stored in the same file. The file also has a section for Wasm binaries, row group metadata, optional sections, a footer, and a fixed-size postscript with metadata size, footer size, compression type, checksum type, data checksum, schema checksum, version, and the F3 magic value.

The format definition says optional metadata sections currently store Wasm binaries, but could also hold things like column UUIDs for schema evolution or zone maps for predicate pushdown.

fff-poc/src/options.rs shows default writer settings: 8 MiB IOUnit size, 64 Ki rows per encoding unit, xxhash checksum, one row group by default, encoder dictionary by default, no per-IOUnit checksum by default, and uncompressed EncUnits by default.

fff-poc/src/writer.rs builds writers around Arrow RecordBatch input, logical column encoders, row group metadata, shared dictionary context, checksum handling, and a Wasm writing context. Comments note that parts of the API are still research-oriented and some custom encoding constraints are not fully general.

fff-poc/src/reader/mod.rs includes a newer FileReaderV2 with projection and selection handling, metadata buffer loading from the postscript/footer, shared dictionary cache, checksum options, and Wasm reading context.

Reproduction and caveats

The documentation includes paper reproduction steps for metadata overhead, compression ratio, decompression speed, random access, dictionary scope, Wasm microbenchmarks, Wasm size, Wasm decoding time, and checksum overhead. Several experiments depend on external forks or code modifications, and one section says a one-off script is not yet provided.

The README says the project was only tested on an Intel machine with Debian 12. Build instructions include submodule initialization, a Debian setup script, cargo build -p fff-poc, and cargo test -p fff-poc.