newsOpenAITrust 88 · LabPublished 1mo agoLive · 1mo ago
Safety and alignment in an era of long-horizon models
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 61%Verisight →
- PossiblePossibly related (embedding) · 58%muxi-ai/muxi →
- PossiblePossibly related (embedding) · 57%opencmit/alphora →
- PossiblePossibly related (embedding) · 56%Harmonizing AI Safety Thresholds →
- PossiblePossibly related (embedding) · 56%MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection →
- PossiblePossibly related (embedding) · 51%Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers →
- PossiblePossibly related (embedding) · 60%Ensuring Safe Physical AI in Urban Mobility via Hazard-Informed Synthesized Envelopes →
- PossiblePossibly related (embedding) · 70%The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems →
Covers
Covers (incoming)
paperRegime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety ClassifierspaperEnsuring Safe Physical AI in Urban Mobility via Hazard-Informed Synthesized EnvelopespaperThe safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systemspaperResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&DpaperOpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party SkillspaperSafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety AlignmentpaperUncensored Open-weight Models: Redistribution as the Persistence LayerpaperCLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety AlignmentpaperWhen Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AIpaperFLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models
Related across the graph
paperUncensored Open-weight Models: Redistribution as the Persistence Layerrepoopencmit/alphorapaperWhen Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AIpaperSafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety AlignmentpaperCLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety AlignmentpaperMedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary InspectionpaperOpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party SkillspaperRegime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety ClassifierspaperThe safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systemspaperFLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Modelsrepomuxi-ai/muxipaperResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&DpaperEnsuring Safe Physical AI in Urban Mobility via Hazard-Informed Synthesized EnvelopespaperHarmonizing AI Safety ThresholdscompanyVerisight
