newsReddit r/MachineLearningTrust 52 · CommunityPublished 4h agoLive · 1h ago

Contrastive Decoding Diffing (CDD): recovering verbatim finetuning data from logits alone, no weight access needed[R]

We built a model diffing method that recovers verbatim content from narrowly finetuned LLMs using only grey-box logit access (no weights, no activations, no probe corpus). Recent work (Minder, Dumas et al., "Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences") showed that finetuning leaves detectable traces in activation differences between base and finetuned models. Their method, Activation Difference Lens (ADL), steers

Covers

paperDistill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation paperLogit-Contribution Scoring Identifies Non-Literal Retrieval Heads paperUnderstanding Evaluation Illusion in Diffusion Large Language Models paperDNA Language Models: An Assessment of Pre-Training for Fine-Tuning Tasks paperOn the Role of Directionality in Structural Generalization

Related across the graph

paperOn the Role of Directionality in Structural Generalization paperUnderstanding Evaluation Illusion in Diffusion Large Language Models paperLogit-Contribution Scoring Identifies Non-Literal Retrieval Heads paperDNA Language Models: An Assessment of Pre-Training for Fine-Tuning Tasks paperDistill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation