34 lines
3.3 KiB
Markdown
34 lines
3.3 KiB
Markdown
# Discord Link Ingest State
|
|
|
|
last_checked_at: 2026-07-18T01:03:50Z
|
|
last_message_created_at: 2026-07-18T00:25:36.573000000Z
|
|
lookback_used: incremental_since_last_message_created_at_after_discrawl_git_share_auto_update
|
|
channels:
|
|
chat: '1028287639918497822'
|
|
tw: '1477793137064935675'
|
|
|
|
## Last run summary - 2026-07-18T01:03:50Z
|
|
|
|
- Discrawl status succeeded. `discrawl --json sql` allowed the configured git share auto-update/import path to run; archive import reported 174,711 share rows and 167,645 messages.
|
|
- Messages scanned: 4 total new #chat/#tw messages newer than `2026-07-17T23:38:31.086000000Z`.
|
|
- URL-containing messages: 1.
|
|
- URLs extracted: 1 mention, 1 normalized unique URL after dedupe.
|
|
- Raw articles saved: 0.
|
|
- Wiki pages created: 0.
|
|
- Existing wiki pages updated: 0.
|
|
- Link-only/skipped: 1 — `https://www.nii.ac.jp/pi/` is the old framed NII “Progress in Informatics” journal homepage; Defuddle returned an empty body and direct frame fetch showed only journal/about/current-issue navigation. It did not meet the strict raw-ingest threshold as a standalone source.
|
|
- Extraction/search blockers: 1 — `web_extract` backend is search-only; fallback direct Python fetch was enough to classify the page as link-only.
|
|
- Interest-profile update: unchanged; this confirms the existing rule that a durable academic/institutional page still needs substantive reusable body content, not just a low-context homepage, before raw/wiki ingest.
|
|
|
|
## Current run processed / notable URLs
|
|
|
|
- https://www.nii.ac.jp/pi/
|
|
|
|
## Dedupe note
|
|
|
|
Dedupe is primarily enforced by scanning `raw/articles` frontmatter `source_url:` values. The state file records recent notable discovery URLs so repeated X/Twitter digest links can be suppressed between runs.
|
|
|
|
## Stable rubric notes
|
|
|
|
Repeated durable interests confirmed so far: LLM Wiki / knowledge-management tooling; autonomous agent loops; verification/evaluator separation; MCP/agent identity security; practical dev-infra; information-integrity and knowledge-governance sources; public/civic infrastructure uses of AI; private/local AI workflows; agent operator observability; agent-oriented CLI design; service-owned agent-readable skill indexes; command-execution bypass research; credential-leakage failure modes; agent-evaluation stacks; implementation-derived quality metrics; graph/HITL/resume workflow orchestration; package supply-chain compromise reports; human-verification advertising; transcript-retention sources when they expose infrastructure tradeoffs; accessibility as an operational capability; AI-safety triage/evaluation frameworks; agent-skill registries with trust boundaries; concrete harness-engineering reliability primitives; AI-crawler economics; request-level agent-payment infrastructure; code-to-repo-wiki maintenance loops; browser-agent harnesses with DOM/network/console/accessibility surfaces; local-government climate-adaptation AI; CI/CD and GitHub Actions security checklists; public secret-leak monitoring; creative-coding/computational-craft and analytics-engineering quality case studies as raw-only watchlist items unless they recur; social-media age assurance when it connects child safety to privacy-preserving identity, platform design responsibility, or KYC risk; public-sector data integration / record linkage sources when they connect data quality, explainability, civic service delivery, and data-protection risks.
|