Safety
50 items across the graph · 2 news stories — tagged with Safety.
Latest news
IL Governor Pritzker Signs Artificial Intelligence Safety Law - WGY
IL Governor Pritzker Signs Artificial Intelligence Safety Law WGY
Read full story →More news · 1
From the graph · 48
UQLM: Uncertainty Quantification for Language Models, is a Python package for UQ-based LLM hallucination detection
Agentlens is a trusted agent trading platform. Here, you can quickly find the Agent that meets your needs, and you can also publish your own Agent to turn it in…
A resource repository for machine unlearning in large language models
Deliver safe & effective language models
The action firewall for AI agents. Enforce policy and human approval before risky tool calls, shell commands, workflows, and production changes, with auditable…
Self-hosted, OpenAI-compatible AI gateway for private RAG, natural-language data access, and tool-calling agents.
Decrypted Generative Model safety files for Apple Intelligence containing filters
Centralized agent control plane for governing runtime agent behavior at scale. Configurable, extensible, and production-ready.
The Execution Security Layer for the Agentic Era. Providing deterministic "Sudo" governance and audit logs for autonomous AI agents.
Agentic AI research papers, benchmarks, frameworks, and tools curated across 24 domains.
Machine learning and Structure-from-Motion tools for estimating vehicle speed from imagery for traffic monitoring, road safety, and autonomous systems research.
A Living Library You Can Talk To. Open-source educational platform with 30 historical figures from philosophy, science, art, mysticism, and activism. Stories, d…
Alignment-research scaffold (autoresearch-style) for LLM guardrails: search over a single policy.md surface
Automated Proof-of-Carrying Change Management for AIOps 2026
Doberman is an AI agent security framework for guardrails, prompt injection defense, runtime policy enforcement, tool-use permissions, agent monitoring, audit l…
An orchestration runtime for multi-agent AI systems. Declare agents, tools, and policies as YAML; Orloj schedules, executes, routes, and governs them for produc…
AgentGuard: Zero-Trust Security Foundation for AI Agents
Open-source adversarial testing engine, SDK, and CLI for AI agents. Runs locally or against the Humanbound Platform.
Template repository with AI agent guardrails, safety protocols, and sprint task framework. For Claude, GPT, Gemini, and all LLMs.
Persistent Claude Code agents with scheduling, sessions, memory, and Telegram.
Claude Code best practices applied to application design. Interactive HLD/LLD visualizations, a DB-governed implementation example, and the same primitives as a…
lintlang is a static linter for AI agent configs, tool descriptions, and system prompts that runs zero-LLM quality gating in CI. Catches language-level failures…
[L0 CONSTITUTION] arifOS — constitutional MCP kernel. Law, identity, F1–F13, VAULT999. Judges but never executes. DITEMPA BUKAN DIBERI.
Biologically-grounded adversarial training platform: cyclic Wake/Dream/Nightmare/Compress phases that accumulate model robustness without catastrophic forgettin…
A curated, continuously updated reading list of 200+ papers on LLM agents: planning, memory, tool use, multi-agent, evaluation & safety. Companion to the survey…
End-to-end pipeline for seeing how LLMs actually process your prompts. Capture attention across every layer, render heatmaps and cooking curves, compare variant…
CORE is a governance runtime for autonomous AI systems. It enforces constitutional rules during execution, prevents governance bypass, and creates auditable aut…
The open-source AI agent control plane: MCP firewall, model gateway with budgets, human approvals, runtime observability, and audit trails
High-fidelity Claude Fable 5 (Mythos) environment emulation and automated multi-agent jailbreak (Pack Hunt) research laboratory.
Native rules, hooks, and guards that prevent Claude Code and Codex from hallucinating code, duplicating files, or shipping unverified changes.
Deterministic guardrails for AI agents — the LLM proposes, your rules dispose. A sub-microsecond, JIT-compiled rule engine in Rust, with a visual Studio.
Lightweight AI safety middleware that protects humans by intercepting self-harm and criminal intent in LLM prompts. Features a 3-stage safety pipeline, MCP serv…
FDE Agent — 把 AI 装进企业的业务流程,离场后 7×24 自己跑。MIT 开源。
A structured workflow system for AI coding agents - harness engineering, execution loop, skills, hooks, and a learning loop. Works with Claude Code, Codex, Anti…
Specs that enforce themselves. Turn specs into contracts that can't be broken by helpful LLMs.
A Socratic supervision layer for AI coding agents.
little-canary is a prompt-injection detector that reads attacks by their effect on a sacrificial canary model before they reach production. Puts a small canary…
A lightweight, highly secure AI API Gateway/Proxy written in Go. Acts as transparent middleware between local AI coding clients (OpenCode/Pi/Cursor) and upstrea…
The Silence of Intelligence — A comprehensive analysis of Anthropic CEO Dario Amodei's philosophy on Scaling Laws, AI safety, and the future of humanity. / Anth…
Open-source EU AI Act compliance scanner. 51 checks across Articles 9-15. Drop-in trust layers for LangChain, CrewAI, AutoGen, OpenAI. Local-first, no data leav…
The specification, reference runtimes, validator, LLM bridge, and conformance suite for URML — an open language for robot intent.
A practical guide to AI privacy, profiling, shadow profiling, local AI, cloud AI, and the future of human autonomy.
Making a model's behavior match human intent and values.
A tiny safety classifier for fast content filtering.
When a model states something fluent but false.
Evidence that model refusals are mediated by one interpretable direction in activation space.
Training models to critique and revise their own outputs against principles.
AI safety and evaluation tooling for production systems.
