HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation
Production LLMs must handle inputs that attempt to override system instructions, bypass safety policies or elicit harmful responses. A common mitigation is a separate guardrail model. Existing reports, however, provide little evidence on Russian prompt injection or Russian surface obfuscation. We present HiveTraceGuard-Pro, a 0.6B generative guardrail LoRA-tuned from Qwen3-0.6B. It is trained on Russian and English and uses one binary scoring rule (safe/unsafe) for the final target turn. Its training corpus pairs harmful examples, where a counterpart exists, with benign examples from the same
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%Open-sourcing a two-stage prompt-injection detector (regex gate + quantised DeBERTa-v3 ONNX), trained partly on real attacks from a game I ran [P] →
- PossiblePossibly related (embedding) · 51%SentryCode: Real-time Auditor + Honeytokens for AI Coding Agents [P] →
- FuzzySimilar title/name (fuzzy) · 84%GoogleCloudPlatform/generative-ai →
“Fuzzy title match (0.92): “HiveTraceGuard-Pro: A Compact Generative Guardrail for Promp” ≈ “GoogleCloudPlatform/generative-ai””
- FuzzySimilar title/name (fuzzy) · 59%NirDiamant/Prompt_Engineering →
“Fuzzy title match (0.73): “HiveTraceGuard-Pro: A Compact Generative Guardrail for Promp” ≈ “NirDiamant/Prompt_Engineering””
- FuzzySimilar title/name (fuzzy) · 59%steven2358/awesome-generative-ai →
“Fuzzy title match (0.73): “HiveTraceGuard-Pro: A Compact Generative Guardrail for Promp” ≈ “steven2358/awesome-generative-ai””
- FuzzySimilar title/name (fuzzy) · 59%linshenkx/prompt-optimizer →
“Fuzzy title match (0.73): “HiveTraceGuard-Pro: A Compact Generative Guardrail for Promp” ≈ “linshenkx/prompt-optimizer””
- LinkedLinked via arxiv author · 85%Nikita Oblakov →
“HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation”
- LinkedLinked via arxiv author · 85%Sabrina Sadiekh →
“HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation”
