feat: support rich content in long-form tweets (X Articles)

Add support for extracting rich content from X's long-form tweets,
including embedded code snippets, markdown blocks, quoted tweets,
and other structured content that was previously lost.

Changes:
- Add fieldToggles with withArticleRichContentState to TweetDetail API request
- Implement Draft.js content_state parser (renderContentState) that converts
  blocks and entities to readable markdown format
- Add content_state type definition to GraphqlTweetResult

Supported content:
- Block types: paragraphs, headers, ordered/unordered lists, blockquotes
- Entity types: MARKDOWN (code blocks), DIVIDER, TWEET, LINK, IMAGE

No impact on regular tweets - rich content only adds payload when present.

Includes 20 unit tests for the parser and an opt-in live smoke test.

Co-authored-by: Christian Catalan <[email protected]>
This commit is contained in:
cc-vps
2026-01-12 05:12:33 +00:00
committed by Peter Steinberger
co-authored by Christian Catalan
parent e0b960db1a
commit 316cdf77be
5 changed files with 615 additions and 0 deletions
+17
View File
@@ -333,4 +333,21 @@ d('live CLI (Twitter/X)', () => {
expect(snapshot.cached).toBe(true);
expect(snapshot.ids && Object.keys(snapshot.ids).length).toBeGreaterThan(0);
});
it('long-form tweet (article) extracts rich content (opt-in)', async () => {
const longformTweetId = (process.env.BIRD_LIVE_LONGFORM_TWEET_ID ?? '').trim();
if (!longformTweetId) {
// Skip unless explicitly provided - long-form tweets may be deleted/unavailable
return;
}
const read = await runBird([...baseArgs, '--cookie-timeout', cookieTimeoutArg, 'read', longformTweetId, '--json'], {
timeoutMs: 45_000,
});
expect(read.exitCode).toBe(0);
const tweet = parseJson<{ id?: string; text?: string }>(read.stdout);
expect(tweet.id).toBe(longformTweetId);
// Long-form tweets (articles) typically have substantial content (>500 chars)
// This verifies the article content is being extracted, not just a stub
expect(tweet.text?.length).toBeGreaterThan(500);
});
});