news-crawler-LM: A Small Long-Context Model For High-Quality News Crawling
Extracting structured content from news pages remains challenging due to heterogeneous HTML layouts, inconsistent markup, and substantial boilerplate such as navigation elements and advertisements. Rule-based news crawlers can achieve high extraction accuracy by encoding site-specific structure, but require manual configuration in order to generalize to new publishers. Large language models provide a more flexible alternative by reducing the need for handcrafted rules, but their high computational cost limits practical deployment. In this paper, we introduce news-crawler-LM, a small long-conte
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 57%Large Language Models Are Still Getting Stronger, but Researchers Face New Bottlenecks in Data, Evaluation, and Safety | Newswise - Newswise →
- PossiblePossibly related (embedding) · 49%Structured PDF-to-JSON: A Guide to Open-Source Extraction Models in 2026 - MarkTechPost →
- PossiblePossibly related (embedding) · 47%Why large language models are an economic dead end - Newsroom →
- LinkedLinked via arxiv author · 85%Pascal Stolzenburg →
“news-crawler-LM: A Small Long-Context Model For High-Quality News Crawling”
- LinkedLinked via arxiv author · 85%Jonas Golde →
“news-crawler-LM: A Small Long-Context Model For High-Quality News Crawling”
- LinkedLinked via arxiv author · 85%Max Dallabetta →
“news-crawler-LM: A Small Long-Context Model For High-Quality News Crawling”
- LinkedLinked via arxiv author · 85%Alan Akbik →
“news-crawler-LM: A Small Long-Context Model For High-Quality News Crawling”
