Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents

Large language model (LLM) agents have shown strong performance in long-horizon tasks that require planning, tool use, and interaction with external environments. However, most existing benchmarks implicitly assume a monolingual setting, where the entire execution process, including reasoning, tool invocation, and output generation, is conducted within a single language. In contrast, real-world applications often involve multilingual inputs and outputs within a unified workflow, yet the interaction between multilinguality and agentic execution remains underexplored. In this work, we introduce

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 63%langwatch/langwatch
  • PossiblePossibly related (embedding) · 62%harvard-cns/orla
  • PossiblePossibly related (embedding) · 62%antoinezambelli/forge
  • PossiblePossibly related (embedding) · 61%run-llama/ParseBench
  • PossiblePossibly related (embedding) · 60%agent-tools
  • PossiblePossibly related (embedding) · 26%langgenius/dify

    Possibly related via embedding similarity 0.57 (not asserted). Timestamp check: artifact slightly before paper (-5d).

  • PossiblePossibly related (embedding) · 27%deepset-ai/haystack

    Possibly related via embedding similarity 0.58 (not asserted). Timestamp check: artifact slightly before paper (-5d).

  • PossiblePossibly related (embedding) · 29%neuml/txtai

    Possibly related via embedding similarity 0.58 (not asserted). Timestamp check: artifact after paper (+6d).

Implements

Related to

Implements (incoming)

Covers (incoming)

authored (incoming)

Related across the graph

Topics