newsReddit r/MachineLearningTrust 52 · CommunityPublished 1mo agoLive · 1mo ago
Best models for generating red-team attacks? Also looking for public datasets [R]
Hi everyone, I'm currently working on a framework to evaluate the security of LLM applications and AI agents, and I've been stuck on one part for a while. Most red-teaming frameworks rely on an LLM to generate adversarial prompts. My question is more about which model to use . Which closed-source models would you recommend for generating high-quality attacks? Which open-source models
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%Online Safety Monitoring for LLMs →
- PossiblePossibly related (embedding) · 45%A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems →
- PossiblePossibly related (embedding) · 45%superlinked/sie →
- PossiblePossibly related (embedding) · 45%PurpleAILAB/Decepticon →
- PossiblePossibly related (embedding) · 59%votal-ai-hq/wb-red-team →
- PossiblePossibly related (embedding) · 53%toxy4ny/redteam-ai-benchmark →
- PossiblePossibly related (embedding) · 59%Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming →
- PossiblePossibly related (embedding) · 58%votal-ai-hq/ai-red-teaming →
Covers
Covers (incoming)
Related across the graph
repovotal-ai-hq/wb-red-teamreposuperlinked/siepaperAn Early Warning of Emerging Biosecurity Risks in Frontier LLMsrepoPurpleAILAB/Decepticonrepotoxy4ny/redteam-ai-benchmarkpaperOnline Safety Monitoring for LLMsrepovotal-ai-hq/ai-red-teamingpaperAgent Hacks Agent: Autoresearch for Production-Agent Red-TeamingpaperA Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems
