repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 3d ago
go-appsec/toolbox
Collaborative application security testing between humans and agents via CLI and MCP
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%PACE: A Proxy for Agentic Capability Evaluation →
- PossiblePossibly related (embedding) · 50%SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions →
- PossiblePossibly related (embedding) · 47%Reasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational study →
- PossiblePossibly related (embedding) · 47%TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution →
- PossiblePossibly related (embedding) · 45%Show HN: macOS data protection keychain for Electron apps →
- PossiblePossibly related (embedding) · 59%Launch HN: Traceforce (YC S26) – Company-wide security monitoring for AI apps →
- PossiblePossibly related (embedding) · 57%Launch HN: Traceforce (YC S26) – Secure AI apps, one device at a time →
Implements
paperPACE: A Proxy for Agentic Capability EvaluationpaperSWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding SessionspaperReasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational studypaperTestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution
Covers (incoming)
Related across the graph
newsLaunch HN: Traceforce (YC S26) – Secure AI apps, one device at a timepaperPACE: A Proxy for Agentic Capability EvaluationnewsLaunch HN: Traceforce (YC S26) – Company-wide security monitoring for AI appspaperReasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational studypaperTestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-EvolutionpaperSWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding SessionsnewsShow HN: macOS data protection keychain for Electron apps
