newsOpenAITrust 88 · LabPublished 4d agoLive · 2d ago
Safety and alignment in an era of long-horizon models
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 61%Verisight →
- PossiblePossibly related (embedding) · 58%muxi-ai/muxi →
- PossiblePossibly related (embedding) · 57%opencmit/alphora →
- PossiblePossibly related (embedding) · 56%Harmonizing AI Safety Thresholds →
- PossiblePossibly related (embedding) · 56%MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection →
- PossiblePossibly related (embedding) · 70%The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems →
- PossiblePossibly related (embedding) · 58%ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D →
- PossiblePossibly related (embedding) · 53%OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills →
Covers
Covers (incoming)
Related across the graph
repoopencmit/alphorapaperMedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary InspectionpaperOpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party SkillspaperThe safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systemsrepomuxi-ai/muxipaperResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&DpaperHarmonizing AI Safety ThresholdscompanyVerisight
