feat: unify research in profile-bound decks with TweetDeck-style UI

This commit is contained in:
2026-09-24 15:59:54 +09:00
parent da8f52e605
commit d2cbf4dbd3
90 changed files with 2617 additions and 6934 deletions
+114 -108
View File
@@ -1,134 +1,140 @@
# Research decks
## Workspace interface
The deck occupies the viewport with a dark sidebar and horizontally arranged,
independently scrolling columns. A compact toolbar names the active deck.
The sidebar switches decks, adds columns, and jumps to a column; on mobile,
it becomes a compact top bar. Column header menus expose editing, ordering,
and deletion. Creation and editing use native modal dialogs with Escape and
focus restoration. Tokens use a navy/blue palette and a system sans font.
This replaces the original spacious Garden reader layout. Layout inspiration:
[Twitter's TweetDeck design notes](https://blog.x.com/en_us/a/2012/designing-the-new-tweetdeck).
## Product direction
A research topic should become a TweetDeck-style workspace: columns represent
questions or perspectives, collect posts across social platforms, and support
AI summaries with references back to the evidence. The intended platforms are
Twitter, Mastodon, Bluesky, Threads, and Nostr.
A research topic becomes a TweetDeck-style workspace: columns represent
questions or perspectives and will eventually collect posts across Twitter,
Mastodon, Bluesky, Threads, and Nostr for AI summaries with source references.
The current implementation supports Twitter only, with manually or
WebMCP-authored decks. Built-in planning, summaries, and other connectors
remain future work.
The first increment is a manually composed deck backed by Twitter. It makes
the deck definition, column lifecycle, and post normalization concrete before
adding AI or more connectors.
## Workspace and column model
## Implemented boundary
Both `/` and `/deck` render the same deck-only application. Separate reader,
search, list, user-profile, and conversation pages have been removed.
Original-post links open X.
- `src/features/decks/model.ts` validates a versioned deck definition: name,
ordered columns with stable IDs, and provider-specific search conditions.
At most six columns are loaded, bounding concurrent first-page requests.
- `src/features/platforms/types.ts` defines the display/evidence record without
importing Bird: stable key, platform, native identity, original URL, text,
author, optional publication time, media, and quoted post.
- `src/features/platforms/twitter.ts` maps existing reader posts into that
record. Keys use `twitter:<id>` rather than the author's mutable handle.
Raw provider responses and credentials do not enter the common record.
- `ResearchColumn` uses the existing Twitter search hook and server functions.
Equal search conditions share the reader's query cache. Distinct conditions
have independent loading/error/pagination state. Profile changes invalidate
queries through the existing profile switcher.
- `ResearchPostCard` only consumes the common record. Provider-specific detail
routes remain in the existing Twitter reader. Original-post links, media,
and one level of quotes are shown in the deck; engagement metrics and
Twitter article previews are not normalized yet.
`src/features/decks/model.ts` validates a version-2 workspace containing
`activeDeckId` and one or more named decks. Each deck has a stable ID and up to
six ordered columns. Only the active deck mounts its columns. The UI supports
creating, selecting, renaming, and deleting deck profiles; the last deck cannot
be deleted. Deleting the active deck selects the first remaining deck.
The current source is a single Twitter search per column. Its `platform`
discriminator is the extension point for another real connector, not a claim
that five adapters already work. Twitter's raw syntax, Top/Latest, and follows
filter are not requirements imposed on the other platforms. Unsupported
source definitions fail validation instead of silently dropping conditions.
Each column has a stable ID, title, required `profileName`, and one source:
The existing reader and its active-feed WebMCP tools remain separate from the
deck. Mounting multiple reader `PostFeed` components would register conflicting
active-feed tools; deck columns therefore use the search hook directly.
On `/deck`, `get_deck` and `set_deck` expose the definition once local storage
has loaded. They unregister when leaving the route.
| Source kind | Conditions |
| --- | --- |
| `search` | Native Twitter `query`, `product` (`Top`/`Latest`), and `following` |
| `user` | `target`: handle or X/Twitter profile URL, normalized to a handle |
| `list` | `target`: numeric ID or X/Twitter list URL, normalized to an ID |
All sources currently require `platform: "twitter"`. Unsupported definitions
fail validation. Column IDs must be unique within a deck, and deck IDs within
the workspace. Manual edits and agent tools use the same final schema.
`profileName` is the relay account binding, distinct from a named deck profile.
The editor discovers names through `/profiles` and lists through the selected
profile. A column's profile can be changed independently. Every feed and list
request carries an explicit profile name; the server confirms it still exists.
Deleted profiles and unavailable discovery produce errors instead of falling
back to another account. There is no browser-wide profile selection.
The complete source and profile participate in query-cache identity. Equal
conditions on the same profile share loaded pages; distinct profiles retain
separate results and cursors. Columns have independent refresh, pagination,
and error state. Pagination and refresh are manual, with no polling.
## Platform boundary
`src/features/platforms/types.ts` defines the display/evidence record without
Bird imports: stable key, platform, native identity, original URL, text,
author, optional publication time, media, and quoted post. The Twitter mapper
uses `twitter:<id>` keys rather than mutable author handles. Raw responses
and credentials do not enter this record. Cards consume the normalized record;
engagement metrics and Twitter article previews are not normalized yet.
The source union is the extension point for future connectors. Twitter search
syntax and ranking controls are provider-specific. When implementing another
connector, add its real schema and server operation, normalize stable identity,
and bind connection details into cache identity. Multiple sources in one
column should wait until a second connector exercises that need; each source
must retain its own opaque continuation and error state.
## Persistence
The workspace is stored under `twitter-lite-research-deck` in localStorage,
including all deck definitions and the active selection. It stores conditions
and relay profile names, not credentials, posts, summaries, or cursors.
Reloading fetches first pages of the selected deck. SSR and the first browser
render show a loading state until storage has been read.
There is no automatic migration of the former single-deck format, which did
not pin profiles to columns. Invalid or older saved data remains untouched
while the UI presents an empty workspace and an error. An explicit saved edit
replaces it. Storage failures are visible: changes still apply in the current
tab, but persistence failures mean they will be lost on reload. Tabs and
devices do not synchronize; the last write to an origin's localStorage wins.
HTTP and HTTPS origins maintain separate workspaces.
## Deck WebMCP tools
`get_deck({})` returns the current definition (including column IDs) and any
storage error. `set_deck` replaces the complete ordered definition and closes
unsaved editor forms. Read before editing, keep IDs of retained columns, and
include every column you want to keep. Omit IDs for new columns. An empty array
clears the deck. The same deck schema validates the entire input before changes.
Both deck routes expose workspace management and active-column reading.
`list_decks` discovers definitions and available relay profiles; `get_deck`
reads a specific or active deck. `set_deck` creates or replaces and activates a
deck; `select_deck` and `delete_deck` operate by ID. `get_column_posts` and
`load_more_column` read or paginate columns in the active deck. See the
[full tool contracts](webmcp-prototype.md).
For example, after discovering a relay profile named `main`, create a deck:
```json
{
"title": "WebMCPの反応",
"columns": [
{ "title": "日本語", "source": { "query": "WebMCP lang:ja" } },
{ "title": "海外の話題", "source": { "query": "WebMCP lang:en", "product": "Top" } }
{
"title": "日本語",
"profileName": "main",
"source": { "kind": "search", "query": "WebMCP lang:ja" }
},
{
"title": "開発者",
"profileName": "main",
"source": { "kind": "user", "target": "@example" }
}
]
}
```
`source.platform` defaults to `twitter`, `product` to `Latest`, and `following`
to false. `set_deck` returns the applied definition and `persisted: true` before
post loading completes. It does not claim that searches succeeded. A storage
failure returns `isError: true` and explains that the in-memory change was
applied but will be lost on reload. Invalid input changes neither UI nor storage.
Omitting `deckId` creates a deck. To edit, read first and include its `deckId`
and every column to retain; keep existing column IDs. Omitted columns are
removed and omitted column IDs are generated. Post loading is asynchronous,
so a successful save does not mean the upstream requests succeeded.
Native WebMCP needs a supported browser and a secure context. Use the HTTPS
Tailscale Serve origin rather than an HTTP tailnet IP. For Vite, allow that exact
hostname through `__VITE_ADDITIONAL_SERVER_ALLOWED_HOSTS` in the dev process's
environment. HTTP and HTTPS origins have separate browser-local deck storage.
## Future AI work
## Persistence and fetching
One deck is saved under `twitter-lite-research-deck` in localStorage. It contains
only conditions and names, not posts, summaries, credentials, or cursors.
SSR and the initial browser render show a loading state before reading it.
Invalid saved data is retained until the user explicitly saves an edit. Storage
read/write failures are surfaced in the UI.
The selected relay profile is browser-wide, not pinned to a column. Reopening
a deck uses that current profile and fetches first pages. Pagination is manual;
there is no polling or background collection. Tabs do not synchronize deck
edits; the last edit written to localStorage wins. LocalStorage is an initial
single-browser workspace, not the eventual research archive.
## Adding the second platform
1. Add a real provider-specific source schema and a validated server-side
search operation. Keep authentication on the server. Extend the source
union and put dispatch at the feed boundary, outside the post card.
2. Implement normalization with stable provider identity. For Mastodon, do not
treat an instance-local numeric ID as globally unique; for Bluesky prefer
the canonical record identity; for Nostr use event identity. The human
original-post URL and the deduplication identity are separate fields.
3. Add a connection reference when per-column accounts, Mastodon instances,
or Nostr relay sets are introduced. Include it in query-cache identity.
4. Support multiple sources within a column only when a second connector can
exercise it. Each source needs its own continuation and error state. Do not
merge opaque upstream cursors into one cross-platform cursor or claim a
globally complete chronological feed from separately ranked searches.
Search capability is connection-dependent. Mastodon documents that status
search depends on the instance's search backend and authentication:
[Mastodon search API](https://docs.joinmastodon.org/methods/search/).
Nostr's full-text search is an optional relay capability:
[NIP-50](https://github.com/nostr-protocol/nips/blob/master/50.md).
Validate actual connection capabilities when those connectors are added.
Threads authorization and Bluesky endpoint behavior must likewise be verified
when implementing those adapters, not inferred from Twitter's contract.
## AI increment
The next product slice should turn a topic into validated column definitions,
then let the user refine them. AI-generated conditions use the same schema as
manual edits. Summaries should reference a persisted collection snapshot
(post keys, original URLs, retrieval time, source/query and profile context),
so a later refresh does not change the evidence behind an earlier claim.
LocalStorage of conditions alone does not provide this evidence store.
ACP or Codex app-server can connect the planner to the application, but they
are not part of the platform data model or this implementation.
A planner can generate definitions through the existing schema and tools.
Summaries will need persisted collection snapshots: post identity and URL,
retrieval time, source conditions, and profile context. Saved conditions alone
do not preserve the evidence behind a summary. ACP or Codex app-server may
connect a future planner, but neither is part of this implementation.
## Verification
Unit tests cover normalization, source/definition validation, ordering and
storage errors. Playwright tests use the existing standalone mock relay through
the real server functions to verify independent pagination, retry, editing,
ordering, reload, deletion/undo, and invalid saved data. They also check viewport
widths 320/375/414/768 and run Axe on the populated deck.
No live SNS search is required by these tests.
Unit tests cover normalization, source/workspace validation, ordering,
profile-specific caching and pagination, profile discovery failures, and
storage behavior. Playwright uses a standalone mock relay through real server
functions to exercise deck switching, column/profile editing, list selection,
pagination, persistence, and native WebMCP. Accessibility checks run on the
workspace. Automated tests do not need a live SNS search.
+81 -96
View File
@@ -1,98 +1,88 @@
# WebMCP prototype
## Scope and design
## Scope and registration
Expose the existing intentional reader to agents through three React-owned
reader tools. The deck additionally exposes `get_deck` and `set_deck` on `/deck`;
see [deck tool contracts](research-decks.md#deck-webmcp-tools).
Search remains available across routes. Feed tools are registered only
while a valid timeline, search, or conversation is open. Returning home or to
an empty search/list page removes the feed tools. Unsupported browsers do not
register tools or fetch posts automatically.
The deck workspace exposes seven React-owned tools on `/` and `/deck`.
`usewebmcp` owns native browser registration and cleanup. There is no polyfill
or external MCP transport. Unsupported browsers retain the manual UI. Tools
are enabled after local storage loads; relay-profile discovery may still be
pending, which `list_decks` reports as `profiles: null`.
`usewebmcp` 5.1.0 owns browser registration and cleanup. No polyfill is
initialized. The root search tool navigates through TanStack Router and awaits
the same query options as the visible feed. Completed data is reused through
`ensureInfiniteQueryData`; an invalidated or in-progress query is awaited
through `fetchInfiniteQuery`, including refreshes after a profile change.
Reading the feed uses its current React Query result. Continuation and the
scroll observer use `fetchNextPage({ cancelRefetch: false })` to join an
existing request. A feed change during continuation returns an error instead
of reporting old results as belonging to the new feed.
The former `search_posts`, `get_loaded_posts`, and `load_more_posts` tools and
standalone reader routes have been removed. Agents manage named decks and
address columns explicitly, including their bound relay profiles.
## Tool contracts
## Workspace tools
### `search_posts`
Search X, show the criteria and results in the UI, and return the first slice
after retrieval. Searches use the currently selected relay profile.
Inputs match the existing search controls:
| Field | Default | Meaning |
| Tool | Input | Behavior |
| --- | --- | --- |
| `q` | empty | Search text, including raw X search syntax |
| `from` | empty | Author handle, optionally prefixed with `@` |
| `since` | empty | Inclusive `YYYY-MM-DD` date using X search semantics |
| `until` | empty | Exclusive `YYYY-MM-DD` date, later than `since` |
| `lang` | `all` | `all`, `ja`, or `en` |
| `content` | `all` | `all`, `images`, `videos`, or `links` |
| `excludeReplies`, `excludeReposts` | `false` | Exclusions |
| `product` | `Latest` | `Latest` or `Top` |
| `following` | `false` | Restrict to followed accounts |
| `list_decks` | `{}` | Return all deck definitions, `activeDeckId`, available `profiles`, and `storageError` |
| `get_deck` | Optional `deckId` | Read a saved deck; omitted ID selects the active deck |
| `set_deck` | Optional `deckId`, required `title` and `columns` | Create when ID is omitted; otherwise replace an existing deck, then activate it |
| `select_deck` | `deckId` | Activate a saved deck and persist the selection |
| `delete_deck` | `deckId` | Permanently remove the definition; cannot delete the last deck |
Supply `q` or `from`. Unknown properties and invalid inputs fail before
navigation. Existing query construction validates dates, handles, and the
512-character compiled-query limit. A concurrent search fails with an
actionable message. Navigating elsewhere during a search prevents a stale
success response.
`set_deck` accepts at most six columns. Each requires `title`, `profileName`
from `list_decks`, and a discriminated `source`. Its `kind` is `search`, `user`,
or `list`; `platform` defaults to `twitter`. Searches require `query`, with
`product` defaulting to `Latest` and `following` to false. User and list sources
require `target` (handle/profile URL or list ID/URL). Unknown source fields,
invalid targets, duplicate column IDs, and unknown profiles fail before saving.
### `get_loaded_posts`
Read before editing. Include every column to keep; omitted columns are removed.
Preserve IDs for retained columns and omit IDs for new ones. An empty columns
array clears a deck. A supplied deck ID must already exist. Successful mutation
closes unsaved editor forms. `set_deck` returns the applied definition,
`persisted: true`, and `posts: "loading-asynchronously"`; searches can fail
independently after the save succeeds.
Return a slice of the current feed without a network request. Accepts
`offset` (default 0, nonnegative integer) and `limit` (default 20, integer
1–50). Conversation results place the selected post first, once.
Deleting the active deck selects the first remaining one. Deletion has no
workspace-tool undo. Storage failures return `isError: true` explaining that
the mutation applied in memory but could not be persisted. Saving replaces
invalid or legacy saved data; there is no automatic legacy migration.
The response includes `status` (`loading`, `ready`, or `error`), `loading`,
and `error` information alongside the common result. This is a snapshot of
all loaded posts, not just posts inside the viewport.
## Column tools
### `load_more_posts`
`get_column_posts` accepts `columnId`, `offset` (default 0, nonnegative integer),
and `limit` (default 20, integer 1–50). It reads already loaded posts without a
network request. Only columns mounted in the active deck are available; select
the deck and allow it to render first.
Accepts `{}`. Wait for one continuation, append it to the UI, and return up to
20 newly appended posts. An in-progress scroll request is shared. Initial
loading or a feed refresh must finish first. A failed continuation can be retried explicitly.
At the end of the feed the result is an empty `posts` array and `hasMore: false`.
`load_more_column` accepts `columnId`. It loads or retries one continuation
using that column's bound profile. Wait for its initial load or refresh before
calling. Concurrent pagination joins the existing request. If the deck or
column changes during the request, the tool reports an error instead of
returning results under the new identity. At the end, it returns no appended
posts and `hasMore: false`.
Both return:
- `column`: ID, title, bound profile, and source definition
- `status`: `loading`, `ready`, or `error`, plus `loading` and `error` details
- `posts`: normalized records with original URLs, identity, text, author,
and available media/quotes
- `loadedCount`, `offset`, and `nextOffset` for slicing deduplicated cached posts
- `hasMore`: whether the current feed has an upstream continuation
Continuation returns up to 20 newly appended posts. Use `nextOffset` with
`get_column_posts` to read additional already loaded records, and
`load_more_column` for an upstream page. These are cache snapshots; manual UI
refreshes and pagination can change the available records.
## Results and errors
Successful tool results contain JSON in an MCP text content block:
Tools return JSON in an MCP text content block. Execution failures set
`isError: true` with `code`, `message`, and `retryable`. Schema failures use
`invalid-input`; other tool failures use `tool-error`. Column snapshots report
underlying relay failures through their `error` field. An empty successful
query is not an error.
- `request`: active feed kind and normalized criteria.
- `posts`: `id`, `author` (username/name), `text`, `textTruncated`, optional
`createdAt`, and the original X `url`.
- `loadedCount`: number of deduplicated, loaded posts.
- `offset`: start of this slice within the loaded feed.
- `nextOffset`: next unread offset within the already loaded feed, or `null`.
- `hasMore`: whether the last loaded page has an upstream continuation.
`list_decks`, `get_deck`, and `get_column_posts` carry `readOnlyHint: true`.
The other tools change local UI or storage. `delete_deck` carries
`destructiveHint: true`; tools returning external posts mark them untrusted.
Annotations are metadata, not authorization controls. No tool writes to X.
Use `get_loaded_posts` with `nextOffset` for loaded posts outside a returned
slice; use `load_more_posts` for an upstream continuation. Scrolling can load
more posts independently, so these fields describe a snapshot. Post text is
capped at 2,000 characters per post and explicitly marked when truncated;
the source URL remains available. Media, quote bodies, and article previews
are not included in this initial text-oriented tool response.
Failures use `isError: true` and JSON containing `code`, `message`, and
`retryable`. Existing relay error details are preserved. A successful empty
search is not an error.
All tools mark returned external content with `untrustedContentHint: true`.
Only `get_loaded_posts` has `readOnlyHint: true`: search and continuation
change local UI state. No tool performs X write actions. Annotations are
metadata, not authorization controls.
## Verification and limits
## Verification and browser setup
```sh
nix develop -c pnpm test
@@ -100,24 +90,19 @@ nix develop -c pnpm typecheck
nix develop -c pnpm test:e2e e2e/integrations/webmcp.test.ts
```
The E2E file enables native Chromium WebMCP/testing flags and uses
`navigator.modelContextTesting` to invoke actual registered tools. The
standalone mock relay supplies deterministic data through the production
server-function boundary. Tests do not contact X or a personal relay.
The E2E tests enable native Chromium WebMCP/testing flags and use
`navigator.modelContextTesting` to invoke actual registered tools. A mock
relay supplies deterministic responses through real server functions.
For interactive testing, enable Chrome's WebMCP testing flag, restart, run
`nix develop -c pnpm dev`, and open the local site with Model Context Tool
Inspector. Registration uses `document.modelContext`; availability is checked
at mount, so reload after changing browser support. No production origin-trial
token or external MCP-client bridge is configured by this prototype.
For interactive testing, enable `chrome://flags/#enable-webmcp-testing`,
restart Chrome, and use Model Context Tool Inspector on the app. Registration
uses `document.modelContext`; reload after changing browser support. Native
WebMCP requires a secure context: local loopback works for development; remote
Tailscale access should use an HTTPS Serve origin. Allow the exact hostname
through `__VITE_ADDITIONAL_SERVER_ALLOWED_HOSTS` in the Vite process environment.
HTTP and HTTPS have separate browser-local workspaces.
The hook does not forward the browser's execution AbortSignal to application
callbacks. Browser cancellation therefore does not guarantee cancellation of
the shared read request. Tools do detect navigation changes before returning
their asynchronous results. Agent task-selection quality still needs manual
evaluation with the consuming agent; deterministic browser tests verify the
tool contracts and UI behavior.
Design references: [Chrome best practices](https://developer.chrome.com/docs/ai/webmcp/best-practices),
[workflow design](https://developer.chrome.com/docs/ai/webmcp/build-tools),
and [usewebmcp](https://github.com/WebMCP-org/npm-packages/tree/main/packages/usewebmcp).
No production origin-trial token or external MCP-client bridge is configured.
Browser cancellation does not guarantee cancellation of a shared feed request.
Agent task-selection quality still needs evaluation with the consuming agent;
automated browser tests verify contracts and UI behavior.