PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents
Large language model (LLM) agents have shown strong performance in long-horizon tasks that require planning, tool use, and interaction with external environments. However, most existing benchmarks implicitly assume a monolingual setting, where the entire execution process, including reasoning, tool invocation, and output generation, is conducted within a single language. In contrast, real-world applications often involve multilingual inputs and outputs within a unified workflow, yet the interaction between multilinguality and agentic execution remains underexplored. In this work, we introduce
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 63%langwatch/langwatch →
- PossiblePossibly related (embedding) · 62%harvard-cns/orla →
- PossiblePossibly related (embedding) · 62%antoinezambelli/forge →
- PossiblePossibly related (embedding) · 61%run-llama/ParseBench →
- PossiblePossibly related (embedding) · 60%agent-tools →
- PossiblePossibly related (embedding) · 26%langgenius/dify →
“Possibly related via embedding similarity 0.57 (not asserted). Timestamp check: artifact slightly before paper (-5d).”
- PossiblePossibly related (embedding) · 27%deepset-ai/haystack →
“Possibly related via embedding similarity 0.58 (not asserted). Timestamp check: artifact slightly before paper (-5d).”
- PossiblePossibly related (embedding) · 29%neuml/txtai →
“Possibly related via embedding similarity 0.58 (not asserted). Timestamp check: artifact after paper (+6d).”
