newsGoogle News — Machine LearningTrust 62 · AggregatorPublished 2mo agoLive · 2mo ago
Understanding Annotator Safety Policy with Interpretability - Apple Machine Learning Research
Understanding Annotator Safety Policy with Interpretability Apple Machine Learning Research
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology →
- PossiblePossibly related (embedding) · 53%Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity →
- PossiblePossibly related (embedding) · 47%Paved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training Regimes →
- PossiblePossibly related (embedding) · 47%Surrogate Fidelity: When Can Open LLMs Explain Closed Ones? →
- PossiblePossibly related (embedding) · 47%Defending Against Harmful Supervision Hidden in Benign Samples →
- PossiblePossibly related (embedding) · 47%Steering Neural Network Training through Interpretable Constraints Based on Partial Dependence →
- PossiblePossibly related (embedding) · 56%interpretml/interpret →
- PossiblePossibly related (embedding) · 46%Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels →
Covers
paperThe Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training MethodologypaperAdversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy AmbiguitypaperPaved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training RegimespaperSurrogate Fidelity: When Can Open LLMs Explain Closed Ones?paperDefending Against Harmful Supervision Hidden in Benign Samples
Covers (incoming)
paperSteering Neural Network Training through Interpretable Constraints Based on Partial Dependencerepointerpretml/interpretpaperSilent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization LevelspaperRegime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety ClassifierspaperWhat Do Compliance Detectors Read? An Audit of Activation Probes and Guard ModelspaperThe Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and RefusalpaperInterpretable Fuzzy Rule-Based Regression Extension for Ex-Fuzzy LibrarypaperWhen Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation
Related across the graph
paperSteering Neural Network Training through Interpretable Constraints Based on Partial Dependencerepointerpretml/interpretpaperWhen Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time ModerationpaperThe Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and RefusalpaperInterpretable Fuzzy Rule-Based Regression Extension for Ex-Fuzzy LibrarypaperAdversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy AmbiguitypaperPaved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training RegimespaperRegime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety ClassifierspaperSilent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization LevelspaperSurrogate Fidelity: When Can Open LLMs Explain Closed Ones?paperDefending Against Harmful Supervision Hidden in Benign SamplespaperThe Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training MethodologypaperWhat Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models
