Topic

Safety

50 items across the graph · 2 news stories — tagged with Safety.

Latest news

NewsGoogle News — AILive · 1mo ago

IL Governor Pritzker Signs Artificial Intelligence Safety Law - WGY

IL Governor Pritzker Signs Artificial Intelligence Safety Law WGY

Read full story →

More news · 1

From the graph · 48

repo
cvs-health/uqlm

UQLM: Uncertainty Quantification for Language Models, is a Python package for UQ-based LLM hallucination detection

repo
ZhangJinHaHaHa/AgentLens

Agentlens is a trusted agent trading platform. Here, you can quickly find the Agent that meets your needs, and you can also publish your own Agent to turn it in…

repo
chrisliu298/awesome-llm-unlearning

A resource repository for machine unlearning in large language models

repo
PacificAI/langtest

Deliver safe & effective language models

repo
cordum-io/cordum

The action firewall for AI agents. Enforce policy and human approval before risky tool calls, shell commands, workflows, and production changes, with auditable…

repo
schmitech/orbit

Self-hosted, OpenAI-compatible AI gateway for private RAG, natural-language data access, and tool-calling agents.

repo
BlueFalconHD/apple_generative_model_safety_decrypted

Decrypted Generative Model safety files for Apple Intelligence containing filters

repo
agentcontrol/agent-control

Centralized agent control plane for governing runtime agent behavior at scale. Configurable, extensible, and production-ready.

repo
node9-ai/node9-proxy

The Execution Security Layer for the Agentic Era. Providing deterministic "Sudo" governance and audit logs for autonomous AI agents.

repo
mahmoudrabie/agentic-ai

Agentic AI research papers, benchmarks, frameworks, and tools curated across 24 domains.

repo
ultralytics/velocity

Machine learning and Structure-from-Motion tools for estimating vehicle speed from imagery for traffic monitoring, road safety, and autonomous systems research.

repo
chipmates/agoracosmica

A Living Library You Can Talk To. Open-source educational platform with 30 historical figures from philosophy, science, art, mysticism, and activism. Stories, d…

repo
SantanderAI/autoguardrails

Alignment-research scaffold (autoresearch-style) for LLM guardrails: search over a single policy.md surface

repo
ChristoAnsek/audited-change-gate

Automated Proof-of-Carrying Change Management for AIOps 2026

repo
fu351/Doberman-Core

Doberman is an AI agent security framework for guardrails, prompt injection defense, runtime policy enforcement, tool-use permissions, agent monitoring, audit l…

repo
OrlojHQ/orloj

An orchestration runtime for multi-agent AI systems. Declare agents, tools, and policies as YAML; Orloj schedules, executes, routes, and governs them for produc…

repo
WhitzardAgent/AgentGuard

AgentGuard: Zero-Trust Security Foundation for AI Agents

repo
humanbound/humanbound

Open-source adversarial testing engine, SDK, and CLI for AI agents. Runs locally or against the Humanbound Platform.

repo
TheArchitectit/agent-guardrails-template

Template repository with AI agent guardrails, safety protocols, and sprint task framework. For Claude, GPT, Gemini, and all LLMs.

repo
JKHeadley/instar

Persistent Claude Code agents with scheduling, sessions, memory, and Telegram.

repo
war851/AI-Governance-Architecture

Claude Code best practices applied to application design. Interactive HLD/LLD visualizations, a DB-governed implementation example, and the same primitives as a…

repo
hermes-labs-ai/lintlang

lintlang is a static linter for AI agent configs, tool descriptions, and system prompts that runs zero-LLM quality gating in CI. Catches language-level failures…

repo
ariffazil/arifos

[L0 CONSTITUTION] arifOS — constitutional MCP kernel. Law, identity, F1–F13, VAULT999. Judges but never executes. DITEMPA BUKAN DIBERI.

repo
Adit-Jain-srm/NightmareNet

Biologically-grounded adversarial training platform: cyclic Wake/Dream/Nightmare/Compress phases that accumulate model robustness without catastrophic forgettin…

repo
js-lee-AI/awesome-llm-agent-papers

A curated, continuously updated reading list of 200+ papers on LLM agents: planning, memory, tool use, multi-agent, evaluation & safety. Companion to the survey…

repo
taylorsatula/TeaLeaves

End-to-end pipeline for seeing how LLMs actually process your prompts. Capture attention across every layer, render heatmaps and cooking curves, compare variant…

repo
DariuszNewecki/CORE

CORE is a governance runtime for autonomous AI systems. It enforces constitutional rules during execution, prevents governance bypass, and creates auditable aut…

repo
preloop/preloop

The open-source AI agent control plane: MCP firewall, model gateway with budgets, human approvals, runtime observability, and audit trails

repo
keirsalterego/jailbreak-fable

High-fidelity Claude Fable 5 (Mythos) environment emulation and automated multi-agent jailbreak (Pack Hunt) research laboratory.

repo
majiayu000/vibeguard

Native rules, hooks, and guards that prevent Claude Code and Codex from hallucinating code, duplicating files, or shipping unverified changes.

repo
Ordo-Engine/Ordo

Deterministic guardrails for AI agents — the LLM proposes, your rules dispose. A sub-microsecond, JIT-compiled rule engine in Rust, with a visual Studio.

repo
Vishisht16/Humane-Proxy

Lightweight AI safety middleware that protects humans by intercepting self-harm and criminal intent in LLM prompts. Features a 3-stage safety pipeline, MCP serv…

repo
KongFangXun/sofagent

FDE Agent — 把 AI 装进企业的业务流程,离场后 7×24 自己跑。MIT 开源。

repo
blundergoat/goat-flow

A structured workflow system for AI coding agents - harness engineering, execution loop, skills, hooks, and a learning loop. Works with Claude Code, Codex, Anti…

repo
Hulupeep/Specflow

Specs that enforce themselves. Turn specs into contracts that can't be broken by helpful LLMs.

repo
Touchpoint-Labs/Gadfly

A Socratic supervision layer for AI coding agents.

repo
hermes-labs-ai/little-canary

little-canary is a prompt-injection detector that reads attacks by their effect on a sacrificial canary model before they reach production. Puts a small canary…

repo
gumieri/nenya

A lightweight, highly secure AI API Gateway/Proxy written in Go. Acts as transparent middleware between local AI coding clients (OpenCode/Pi/Cursor) and upstrea…

repo
Leading-AI-IO/the-silence-of-intelligence

The Silence of Intelligence — A comprehensive analysis of Anthropic CEO Dario Amodei's philosophy on Scaling Laws, AI safety, and the future of humanity. / Anth…

repo
airblackbox/airblackbox

Open-source EU AI Act compliance scanner. 51 checks across Articles 9-15. Drop-in trust layers for LangChain, CrewAI, AutoGen, OpenAI. Local-first, no data leav…

repo
URML-MARS/URML

The specification, reference runtimes, validator, LLM bridge, and conformance suite for URML — an open language for robot intent.

repo
cnaebadi/ai-disclosure-handbook

A practical guide to AI privacy, profiling, shadow profiling, local AI, cloud AI, and the future of human autonomy.

glossary term
Alignment

Making a model's behavior match human intent and values.

model
Nano-Refuse-0.4B

A tiny safety classifier for fast content filtering.

glossary term
Hallucination

When a model states something fluent but false.

paper
Refusal as a single linear direction

Evidence that model refusals are mediated by one interpretable direction in activation space.

paper
Constitutional methods for alignment

Training models to critique and revise their own outputs against principles.

company
Verisight

AI safety and evaluation tooling for production systems.

Related topics