newsReddit r/MachineLearningTrust 52 · CommunityPublished 1mo agoLive · 1mo ago
What does "Safe AI" look like? [D]
For open-weight LLMs, how practical is it to study defenses against post-release fine-tuning that weakens refusal or safety behavior? I've been seeing “uncensored” or “heretic” variants of new models appear very quickly after release, which raises a question I’m curious about: is fine-tuning resistance a meaningful safety goal for open-weight releases, or is it too narrow because determined users can always modify weights, switch models, o
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%The case for open weights →
- PossiblePossibly related (embedding) · 52%EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures →
- PossiblePossibly related (embedding) · 50%Online Safety Monitoring for LLMs →
- PossiblePossibly related (embedding) · 49%Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring →
- PossiblePossibly related (embedding) · 49%Verisight →
- PossiblePossibly related (embedding) · 48%Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle Vulnerabilities →
- PossiblePossibly related (embedding) · 50%An Early Warning of Emerging Biosecurity Risks in Frontier LLMs →
- PossiblePossibly related (embedding) · 48%Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs →
Covers
Covers (incoming)
paperEvaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle VulnerabilitiespaperAn Early Warning of Emerging Biosecurity Risks in Frontier LLMspaperJudge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMspaperOpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills
Related across the graph
paperBehind the Refusal: Determining Guardrail Activation via Behavioral MonitoringpaperJudge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMspaperAn Early Warning of Emerging Biosecurity Risks in Frontier LLMspaperOpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party SkillsarticleThe case for open weightspaperEvaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle VulnerabilitiespaperEvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety FailurespaperOnline Safety Monitoring for LLMscompanyVerisight
