add
This commit is contained in:
@@ -1,13 +1,79 @@
|
||||
# Discord Link Ingest State
|
||||
|
||||
last_checked_at: 2026-06-28T11:58:52Z
|
||||
last_checked_at: 2026-06-29T16:18:54Z
|
||||
last_message_created_at: 2026-06-28T07:26:20.859000000Z
|
||||
lookback_used: 24h_initial_state_missing
|
||||
lookback_used: incremental_since_last_message_created_at
|
||||
channels:
|
||||
chat: '1028287639918497822'
|
||||
tw: '1477793137064935675'
|
||||
|
||||
## Last run summary — 2026-06-28
|
||||
## Last run summary — 2026-06-29T16:18:54Z
|
||||
|
||||
- Messages scanned: 0 new messages after `2026-06-28T07:26:20.859000000Z` in monitored channels.
|
||||
- URLs found: 0
|
||||
- Raw articles saved: 0
|
||||
- Wiki pages created/updated: 0
|
||||
- Discrawl status: archive current; last_sync_at `2026-06-28T07:46:17Z`; share repo needs update but read-only ingest did not mutate it.
|
||||
- Rubric note: no new evidence; interest profile unchanged.
|
||||
|
||||
## Previous run summary — 2026-06-29T15:16:46Z
|
||||
|
||||
- Messages scanned: 0 new messages after `2026-06-28T07:26:20.859000000Z` in monitored channels.
|
||||
- URLs found: 0
|
||||
- Raw articles saved: 0
|
||||
- Wiki pages created/updated: 0
|
||||
- Discrawl status: archive current; last_sync_at `2026-06-28T07:46:17Z`; share repo needs update but read-only ingest did not mutate it.
|
||||
- Rubric note: no new evidence; interest profile unchanged.
|
||||
|
||||
## Previous run summary — 2026-06-29T14:14:32Z
|
||||
|
||||
- Messages scanned: 0 new messages after `2026-06-28T07:26:20.859000000Z` in monitored channels.
|
||||
- URLs found: 0
|
||||
- Raw articles saved: 0
|
||||
- Wiki pages created/updated: 0
|
||||
- Discrawl status: archive current; last_sync_at `2026-06-28T07:46:17Z`; share repo needs update but read-only ingest did not mutate it.
|
||||
- Rubric note: no new evidence; interest profile unchanged.
|
||||
|
||||
## Previous run summary — 2026-06-29T13:12:27Z
|
||||
|
||||
- Messages scanned: 0 new messages after `2026-06-28T07:26:20.859000000Z` in monitored channels.
|
||||
- URLs found: 0
|
||||
- Raw articles saved: 0
|
||||
- Wiki pages created/updated: 0
|
||||
- Discrawl status: archive current; last_sync_at `2026-06-28T07:46:17Z`; share repo needs update but read-only ingest did not mutate it.
|
||||
- Rubric note: no new evidence; interest profile unchanged.
|
||||
|
||||
## Manual ingest summary — 2026-06-29
|
||||
|
||||
- Trigger: current Discord reply, `Ingest リバエン`, source `https://github.com/bethington/ghidra-mcp`.
|
||||
- Raw articles saved: 1 (`raw/articles/ghidra-mcp-2026.md`)
|
||||
- Wiki pages created: 2 (`entities/ghidra-mcp.md`, `concepts/ai-assisted-reverse-engineering.md`)
|
||||
- Wiki pages updated: 1 (`index.md`)
|
||||
- Rubric note: reverse-engineering/security developer tools with concrete MCP/agent workflows should score high when they show reusable practice, quality enforcement, or unusual automation patterns.
|
||||
|
||||
## Previous run summary — 2026-06-29T12:10:26Z
|
||||
|
||||
- Messages scanned: 0 new messages after `2026-06-28T07:26:20.859000000Z` in monitored channels.
|
||||
- URLs found: 0
|
||||
- Raw articles saved: 0
|
||||
- Wiki pages created/updated: 0
|
||||
- Discrawl status: archive current; last_sync_at `2026-06-28T07:46:17Z`; share repo needs update but read-only ingest did not mutate it.
|
||||
- Rubric note: no new evidence; interest profile unchanged.
|
||||
|
||||
## Previous run summary — 2026-06-29T11:08:07Z
|
||||
|
||||
- Messages scanned: 0 new messages after `2026-06-28T07:26:20.859000000Z` in monitored channels.
|
||||
- URLs found: 0
|
||||
- Raw articles saved: 0
|
||||
- Wiki pages created/updated: 0
|
||||
- Discrawl status: archive current; last_sync_at `2026-06-28T07:46:17Z`; share repo needs update but read-only ingest did not mutate it.
|
||||
- Rubric note: no new evidence; interest profile unchanged.
|
||||
|
||||
## Earlier run summary — 2026-06-29T09:03:18Z to 2026-06-28T14:15:26Z
|
||||
|
||||
Repeated hourly checks found no new messages after `2026-06-28T07:26:20.859000000Z`; no URLs, raw sources, wiki updates, or rubric changes.
|
||||
|
||||
## Last ingest summary — 2026-06-28
|
||||
|
||||
- Messages scanned: 110
|
||||
- URL-containing messages: 103
|
||||
@@ -18,17 +84,6 @@ channels:
|
||||
- Wiki pages updated: 3 (`concepts/llm-wiki-pattern.md`, `concepts/wiki-maintenance-loop.md`, `index.md`)
|
||||
- Link-only / failed extraction: Obsidian Headless help (`defuddle` returned empty), alphaxiv 2606.25331 (`defuddle` returned empty), many X/t.co/media/news links below strict threshold.
|
||||
|
||||
## Raw articles saved this run
|
||||
|
||||
- https://github.com/nashsu/llm_wiki → raw/articles/nashsu-llm-wiki-2026.md (score 4)
|
||||
- https://hermes-agent.nousresearch.com/docs/user-guide/skills/bundled/research/research-llm-wiki → raw/articles/hermes-research-llm-wiki-skill-2026.md (score 4)
|
||||
- https://github.com/emacsmirror/howm → raw/articles/howm-2026.md (score 2)
|
||||
- https://github.com/Imbad0202/academic-research-skills → raw/articles/academic-research-skills-2026.md (score 2)
|
||||
- https://refactoringenglish.com/excerpts/write-an-effective-design-doc/ → raw/articles/refactoring-english-effective-design-doc-2026.md (score 2)
|
||||
- https://github.blog/developer-skills/github/i-automated-my-job-and-it-made-me-a-better-leader/ → raw/articles/github-automated-my-job-better-leader-2026.md (score 2)
|
||||
- https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens-in → raw/articles/chinatalk-transfer-station-economy-2026.md (score 2)
|
||||
- https://github.com/BRO3886/rem → raw/articles/rem-cli-macos-reminders-2026.md (score 2)
|
||||
|
||||
## Processed high-signal URLs
|
||||
|
||||
- https://github.com/nashsu/llm_wiki
|
||||
@@ -44,4 +99,4 @@ channels:
|
||||
|
||||
## Rubric note
|
||||
|
||||
First run confirmed unusually dense interest in LLM Wiki / knowledge-management tooling (`llm_wiki`, Hermes bundled skill docs, howm, Obsidian Headless) from #chat. Keep scoring this cluster high, but continue strict wiki-page updates: create/update pages only for sources that directly change the LLM Wiki operating model; save adjacent workflow/security/dev-tool links as raw-only unless they connect to existing pages.
|
||||
First ingest confirmed unusually dense interest in LLM Wiki / knowledge-management tooling (`llm_wiki`, Hermes bundled skill docs, howm, Obsidian Headless) from #chat. Keep scoring this cluster high, but continue strict wiki-page updates: create/update pages only for sources that directly change the LLM Wiki operating model; save adjacent workflow/security/dev-tool links as raw-only unless they connect to existing pages.
|
||||
|
||||
@@ -47,7 +47,7 @@ sha256: <hex digest of raw body>
|
||||
## Tag Taxonomy
|
||||
- AI and automation: llm, agent, automation, evaluation
|
||||
- Knowledge systems: wiki, knowledge-base, retrieval, synthesis, maintenance, obsidian, markdown
|
||||
- Development practice: dev-tool, cli, git, workflow, quality, security, supply-chain, reliability
|
||||
- Development practice: dev-tool, cli, git, workflow, quality, security, supply-chain, reliability, data-format
|
||||
- Design and hacks: design, hack, interface, hn-like
|
||||
- Public interest: public-interest, civic-tech, accessibility, inclusive-design, media, disinformation, information-integrity, data-protection, privacy, freedom-expression, law
|
||||
- Entities and roles: person, organization, tool, human-in-the-loop
|
||||
|
||||
@@ -0,0 +1,29 @@
|
||||
---
|
||||
title: AI-Assisted Reverse Engineering
|
||||
created: 2026-06-29
|
||||
updated: 2026-06-29
|
||||
type: concept
|
||||
tags: [security, dev-tool, agent, automation, quality]
|
||||
sources: [raw/articles/ghidra-mcp-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
# AI-Assisted Reverse Engineering
|
||||
|
||||
AI-assisted reverse engineering は、binary 解析、逆コンパイル、型復元、関数名付け、コメント付け、デバッグ観察を、人間だけの手作業ではなく AI エージェントと専門道具の組で進める考え方。[[ghidra-mcp]] はこの形を Ghidra と MCP で具体化している。
|
||||
|
||||
重要なのは、LLM に「それらしく読ませる」だけでは品質が安定しないこと。逆解析では、関数名、型、構造体、呼び出し関係、根拠コメントが後続作業の足場になるため、一度の推測ミスが広く伝播する。Ghidra MCP の README は、命名規則、型変更の拒否、文書化の完全性得点、batch operation、transaction といった仕組みを通じて、作業のばらつきを道具側で抑えようとしている。
|
||||
|
||||
## 見るべき軸
|
||||
|
||||
- **読み取りから書き込みへ**: AI が decompile 結果を要約するだけでなく、rename、retype、comment、structure creation まで行うなら、取り消しや検査の境界が必要になる。
|
||||
- **作法の固定**: 命名や型付けの規則を prompt ではなく tool layer に置くと、[[ai-research-automation]] と同じく、人間が毎回注意しなくても運用が続く。
|
||||
- **静的解析と動的観察の接続**: P-code 実行や debugger 連携は、単なる text reading ではなく、仮説を実行で確かめる経路を作る。
|
||||
- **共同作業と版違い**: Ghidra Server、version control、function hash による文書化移植は、複数人・複数版 binary の解析で効く。
|
||||
|
||||
## Open Questions
|
||||
|
||||
- AI が書き込んだ rename/type/comment を、どの段階で人間が承認するべきか。
|
||||
- convention enforcement が強すぎると、未知の binary に必要な例外を潰さないか。
|
||||
- MCP 経由の Ghidra 操作ログを、[[wiki-maintenance-loop]] の raw source や監査記録としてどう残すか。
|
||||
- 逆解析支援道具を扱うとき、研究・防御・教育の用途と悪用可能性の線引きをどう説明するか。
|
||||
@@ -0,0 +1,28 @@
|
||||
---
|
||||
title: Extensible Data File Formats
|
||||
created: 2026-06-29
|
||||
updated: 2026-06-29
|
||||
type: concept
|
||||
tags: [data-format, dev-tool, reliability, supply-chain]
|
||||
sources: [raw/articles/f3-file-format-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
# Extensible Data File Formats
|
||||
|
||||
Extensible data file formats は、保存した時点の読み書き実装だけに縛られず、あとから新しい符号化、圧縮、索引、読み出し単位を足せるようにしたデータ形式の考え方。[[f3-file-format]] はこの方向を、列指向分析ファイルと WebAssembly 復号器の組み合わせで試している。
|
||||
|
||||
既存の広く使われる形式は、互換性が強みである一方、一度広まった配置や符号化の前提を変えにくい。F3 の主張は、ファイルを「固定仕様に従うバイト列」としてだけでなく、「自分を読むためのメタデータと小さな実行物を持つ保存物」として設計すれば、古い読者と新しい符号化のあいだに橋をかけられる、というもの。
|
||||
|
||||
## 見るべき軸
|
||||
|
||||
- **自己説明性**: schema、列メタデータ、checksum、辞書、任意メタデータが、読み手の推測ではなくファイルに残るか。
|
||||
- **実行可能な拡張**: 未対応の符号化を Wasm などで安全に復号できるか。
|
||||
- **保存と実行の境界**: ファイルに実行物を含めると、互換性は増えるが、検証、砂箱化、供給網、安全性の設計が必要になる。
|
||||
- **再現性**: ベンチや論文の主張が、どの入力、どの環境、どの外部 fork に依存しているかを追えるか。
|
||||
|
||||
## 他のページとの関係
|
||||
|
||||
[[ghidra-mcp]] は、専門作業の手順や品質基準を道具の入口に寄せる例。extensible data file formats は、データを読むための知識をファイル形式側に寄せる例。どちらも [[wiki-maintenance-loop]] と同じく、後から読む人・使う人が毎回文脈を再発見しなくてよい状態を作ろうとしている。
|
||||
|
||||
[[llm-wiki-pattern]] との共通点は、情報をただ置くのではなく、将来の読み手が利用できる形に編み直すこと。ただし LLM Wiki は人間と LLM のための知識整理であり、F3 は分析システムのためのバイナリ保存形式である。
|
||||
@@ -0,0 +1,32 @@
|
||||
---
|
||||
title: F3 File Format
|
||||
created: 2026-06-29
|
||||
updated: 2026-06-29
|
||||
type: entity
|
||||
tags: [tool, dev-tool, data-format, reliability, hack]
|
||||
sources: [raw/articles/f3-file-format-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
# F3 File Format
|
||||
|
||||
F3 は、分析用データを保存するための新しい列指向ファイル形式を試す研究実装。Parquet や ORC のような既存形式が長く使われるなかで、古い配置や拡張しにくさを抱え続ける問題に対し、ファイルの中にデータ、メタデータ、必要なら WebAssembly の復号器まで入れる設計を提案している。
|
||||
|
||||
中心にある発想は [[extensible-data-file-formats]]。読み手が符号化方式をすでに知っていればそのまま読める。知らない場合でも、ファイル内の Wasm 復号器を使って読めるようにする。つまり「形式を読むための知識」を、実装や外部仕様だけに置かず、ファイル自身にも持たせる。
|
||||
|
||||
## 何が面白いか
|
||||
|
||||
- **将来対応の逃げ道**: 新しい圧縮・符号化・配置を試すたびにファイル形式全体を作り直すのではなく、符号化単位ごとに復号器を付ける道を作る。
|
||||
- **細かい読み出しの設計**: row group、IOUnit/Chunk、EncUnit という階層で、入出力単位と符号化単位を分けて扱う。
|
||||
- **自己説明性**: Arrow schema、列メタデータ、共有辞書、checksum、Wasm binary、任意メタデータを同じファイルにまとめる。
|
||||
- **研究としての検証**: `fff-bench` と再現手順があり、Parquet、Vortex、Lance、Nimble、ORC などとの比較を論文側で扱っている。
|
||||
|
||||
## 注意点
|
||||
|
||||
README は明確に「研究プロトタイプであり本番利用しない」と書いている。実装にも、研究用 API、未整備の再現手順、Debian 12/Intel でのみ試験済みという制約が残っている。現時点では採用候補というより、データ形式の設計案と実験台として読むのがよい。
|
||||
|
||||
## Wiki 上の位置づけ
|
||||
|
||||
F3 は、[[ghidra-mcp]] が専門道具の操作を MCP の道具層へ寄せたのと少し似ている。どちらも、知識や手順を人間の記憶や prompt だけに置かず、扱う対象の近くに実行可能な形で置こうとする。ただし F3 の対象は逆解析ではなく、長く残るデータファイルそのもの。
|
||||
|
||||
また [[llm-wiki-pattern]] が raw source をあとから読める形で残し、合成された知識と出典を分けるのに対し、F3 はデータと読むための手掛かりを同じ保存物に近づける。どちらも「将来の読み手が困らないように、文脈を一緒に残す」設計として見られる。
|
||||
@@ -0,0 +1,31 @@
|
||||
---
|
||||
title: Ghidra MCP
|
||||
created: 2026-06-29
|
||||
updated: 2026-06-29
|
||||
type: entity
|
||||
tags: [tool, dev-tool, security, agent, automation]
|
||||
sources: [raw/articles/ghidra-mcp-2026.md]
|
||||
confidence: medium
|
||||
---
|
||||
|
||||
# Ghidra MCP
|
||||
|
||||
Ghidra MCP は、Ghidra の静的解析・逆コンパイル・デバッグ機能を、Model Context Protocol 経由で AI エージェントから扱えるようにする Ghidra 拡張と MCP サーバー。README は「デモ用の読み取りだけ」ではなく、関数名変更、型付け、コメント、構造体作成、スクリプト実行、P-code 実行、実機デバッガー連携まで含む 251 個の道具を掲げている。
|
||||
|
||||
このリポジトリが面白いのは、[[ai-assisted-reverse-engineering]] を「AI が逆コンパイル結果を読む」だけで終わらせず、逆解析作業そのものの作法を道具側に寄せている点。v5.0 では命名規則や型の安全性、文書化基準を MCP の道具層で検査し、自動修正・警告・拒否に分ける。これにより、モデルや作業者が変わっても、関数名・型・コメントの粒度がばらけにくい。
|
||||
|
||||
## 目立つ機能
|
||||
|
||||
- Ghidra の関数解析、呼び出し関係、参照、メモリ読み取り、文字列探索、import/export 解析を MCP から呼び出せる。
|
||||
- P-code グラフを使った値伝播、関数単体の P-code 実行、API hash 解決など、静的解析と軽い実行をつなぐ機能がある。
|
||||
- Ghidra TraceRmi による実機デバッガー連携で、breakpoint、register、memory、step、静的アドレスと動的アドレスの対応を扱う。
|
||||
- 関数 hash によって、別版の binary に文書化を移す cross-binary documentation transfer を掲げている。
|
||||
- GUI 付き Ghidra だけでなく、headless server や Docker/CI 向けの運用も想定している。
|
||||
|
||||
## Wiki 上の位置づけ
|
||||
|
||||
[[ai-research-automation]] や [[wiki-maintenance-loop]] が「情報収集・知識化の手順を持続させる」話だとすると、Ghidra MCP は「専門道具の操作手順と品質基準を MCP の道具として固定する」例。LLM に自由入力で作法を毎回思い出させるのではなく、操作の入口が規則を知っている形にする。
|
||||
|
||||
逆解析は security と dev-tool の境界にあるため、このページでは攻撃手順ではなく、解析支援、再現性、共同作業、道具側の品質保証という観点で扱う。
|
||||
|
||||
F3 の [[f3-file-format]] とは領域が違うが、「使い方の知識を対象の近くに置く」という点では響き合う。Ghidra MCP は解析手順を道具層に置き、F3 はデータを読むための復号手段をファイル側に近づける。
|
||||
@@ -2,11 +2,13 @@
|
||||
|
||||
> Content catalog. Every wiki page listed under its type with a one-line summary.
|
||||
> Read this first to find relevant pages for any query.
|
||||
> Last updated: 2026-06-28 | Total pages: 17
|
||||
> Last updated: 2026-06-29 | Total pages: 21
|
||||
|
||||
## Entities
|
||||
|
||||
- [[f3-file-format]] — WebAssembly 復号器をファイル内に同梱し、将来の符号化にも対応しようとする列指向データファイル形式の研究実装。
|
||||
- [[david-erdos]] — データ保護、プライバシー、表現・報道・研究の自由の均衡を研究する Cambridge 法学者。
|
||||
- [[ghidra-mcp]] — Ghidra の逆解析機能を MCP 経由で AI エージェントから扱うための拡張とサーバー。
|
||||
- [[llm-wiki-app]] — Karpathy の LLM Wiki pattern を desktop app、queue、graph/search、MCP/API 付きで具体化する実装。
|
||||
- [[maria-ressa]] — Rappler 共同創業者・ノーベル平和賞受賞者。SNS 上の偽情報、情報操作、報道機関への攻撃を公共性の観点から扱う。
|
||||
- [[obsidian]] — LLM Wiki を閲覧・編集するための Markdown/リンク対応ノートアプリ。
|
||||
@@ -15,10 +17,12 @@
|
||||
|
||||
## Concepts
|
||||
|
||||
- [[ai-assisted-reverse-engineering]] — Ghidra などの専門道具と AI エージェントを組み合わせ、逆解析の命名・型付け・文書化を支援する考え方。
|
||||
- [[ai-research-automation]] — 検索語、巡回先、Slack の場所、人を自動で見直しながら、AI 関連情報を継続収集して報告にまとめる運用。
|
||||
- [[ai-developer-liability]] — AI の出力や悪用だけでなく、開発者のシステム設計そのものにどこまで責任を問えるかという論点。
|
||||
- [[data-protection-and-expression]] — データ保護と、報道・研究・表現の自由が衝突する場面の均衡を扱う論点。
|
||||
- [[digital-gardening-cms]] — メモ、Wiki、作品集、公開サイトを統合し、分類よりリンクと永続性を重視する CMS 設計案。
|
||||
- [[extensible-data-file-formats]] — データ形式が新しい符号化・圧縮・読み出し方を後から受け入れられるようにする設計思想。
|
||||
- [[inclusive-design]] — 外から見えにくい困難や違いを、本人が毎回説明しなくても周囲の配慮につなげる設計。
|
||||
- [[information-integrity]] — 情報操作、偽情報、報道、情報基盤の責任を、公共圏の品質として扱う概念。
|
||||
- [[llm-wiki-pattern]] — LLM が raw source を読み、持続的な相互リンク付き Markdown wiki にコンパイルする運用パターン。
|
||||
|
||||
@@ -93,3 +93,20 @@
|
||||
- Created: concepts/ai-research-automation.md
|
||||
- Updated: concepts/wiki-maintenance-loop.md
|
||||
- Updated: index.md
|
||||
|
||||
## [2026-06-29] ingest | Ghidra MCP
|
||||
- Source saved: raw/articles/ghidra-mcp-2026.md
|
||||
- Created: entities/ghidra-mcp.md
|
||||
- Created: concepts/ai-assisted-reverse-engineering.md
|
||||
- Updated: index.md
|
||||
- Discord context: `Ingest リバエン`; archive lookup did not yet include the current message, so raw provenance records the current chat context without a Discord message id.
|
||||
|
||||
## [2026-06-29] ingest | F3 future-proof file format
|
||||
- Source saved: raw/articles/f3-file-format-2026.md
|
||||
- Created: entities/f3-file-format.md
|
||||
- Created: concepts/extensible-data-file-formats.md
|
||||
- Updated: entities/ghidra-mcp.md
|
||||
- Updated: SCHEMA.md
|
||||
- Updated: index.md
|
||||
- Discord context: repository link shared with `tldr`, followed by `Ingest`.
|
||||
|
||||
|
||||
@@ -0,0 +1,54 @@
|
||||
---
|
||||
source_url: https://github.com/future-file-format/F3
|
||||
ingested: 2026-06-29
|
||||
sha256: 7e4e455b1bdb0733467b76d96e3d27f23857716766b9093c7922de9a08c23b1a
|
||||
---
|
||||
# F3: The Open-Source Data File Format for the Future
|
||||
|
||||
Source: https://github.com/future-file-format/F3
|
||||
Homepage / paper DOI: https://dl.acm.org/doi/10.1145/3749163
|
||||
Repository inspected: future-file-format/F3, commit bd92506 on main
|
||||
GitHub metadata observed on 2026-06-29: Rust, MIT license, 727 stars, 25 forks, 1 open issue, not archived.
|
||||
|
||||
## Repository README summary
|
||||
|
||||
F3 is a data file format designed for efficiency, interoperability, and extensibility. The README frames it as a way to fix layout shortcomings of last-generation formats such as Parquet while preserving future-proof interoperability through embedded Wasm decoders.
|
||||
|
||||
The repository explicitly says this is a research prototype for verifying the paper's ideas and should not be used in production.
|
||||
|
||||
Important directories listed by the project:
|
||||
|
||||
- `format`: FlatBuffer definition of the file format.
|
||||
- `fff-poc`: main code for the F3 format, referencing `fff-core`, `fff-encoding`, `fff-format`, and `fff-ude-wasm`.
|
||||
- `fff-bench`: benchmarks and experiments from the paper.
|
||||
- `fff-ude*`: User-Defined-Encoding and Wasm decoding implementation.
|
||||
- `scripts` and `exp_scripts`: experiment scripts.
|
||||
|
||||
## Implementation notes from local inspection
|
||||
|
||||
The Rust workspace contains the core proof of concept, encoding support, FlatBuffers format generation, User-Defined-Encoding support, Wasm support, benchmark code, and sample Wasm libraries. Rough source size from the inspected tree:
|
||||
|
||||
- `fff-poc`: 39 files, 9,390 lines.
|
||||
- `fff-core`: 6 files, 917 lines.
|
||||
- `fff-encoding`: 10 files, 1,652 lines.
|
||||
- `fff-format`: 3 files, 87 lines.
|
||||
- `fff-ude`: 3 files, 651 lines.
|
||||
- `fff-ude-wasm`: 7 files, 1,735 lines.
|
||||
- `fff-bench`: 46 files, 7,454 lines.
|
||||
- `wasm-libs`: 22 files, 1,351 lines.
|
||||
|
||||
`format/File.fbs` describes the file layout as row groups containing IOUnits/Chunks, which contain EncUnits. EncUnits are contiguous and can carry encoding metadata; an encoding may point to a Wasm binary stored in the same file. The file also has a section for Wasm binaries, row group metadata, optional sections, a footer, and a fixed-size postscript with metadata size, footer size, compression type, checksum type, data checksum, schema checksum, version, and the `F3` magic value.
|
||||
|
||||
The format definition says optional metadata sections currently store Wasm binaries, but could also hold things like column UUIDs for schema evolution or zone maps for predicate pushdown.
|
||||
|
||||
`fff-poc/src/options.rs` shows default writer settings: 8 MiB IOUnit size, 64 Ki rows per encoding unit, xxhash checksum, one row group by default, encoder dictionary by default, no per-IOUnit checksum by default, and uncompressed EncUnits by default.
|
||||
|
||||
`fff-poc/src/writer.rs` builds writers around Arrow `RecordBatch` input, logical column encoders, row group metadata, shared dictionary context, checksum handling, and a Wasm writing context. Comments note that parts of the API are still research-oriented and some custom encoding constraints are not fully general.
|
||||
|
||||
`fff-poc/src/reader/mod.rs` includes a newer `FileReaderV2` with projection and selection handling, metadata buffer loading from the postscript/footer, shared dictionary cache, checksum options, and Wasm reading context.
|
||||
|
||||
## Reproduction and caveats
|
||||
|
||||
The documentation includes paper reproduction steps for metadata overhead, compression ratio, decompression speed, random access, dictionary scope, Wasm microbenchmarks, Wasm size, Wasm decoding time, and checksum overhead. Several experiments depend on external forks or code modifications, and one section says a one-off script is not yet provided.
|
||||
|
||||
The README says the project was only tested on an Intel machine with Debian 12. Build instructions include submodule initialization, a Debian setup script, `cargo build -p fff-poc`, and `cargo test -p fff-poc`.
|
||||
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user