Infrastructure
50 items across the graph — tagged with Infrastructure.
From the graph · 50
Graph-Native Infrastructure for Context and Accountable AI Systems
Monitoring for Proxmox, Docker, Kubernetes, TrueNAS, and vSphere that watches your infrastructure for you: smart alerts, AI patrols that catch silent failures,…
Democratizing Reinforcement Learning for LLMs
Local persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.
Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-groun…
AI Infrastructure Engineer Learning Track - Production ML infrastructure curriculum (2-4 years experience)
Open-source, self-hosted AI app builder — an agent builds real apps in isolated sandboxes on your own server, each live at a preview URL. Self-host in one comma…
Flexible Fullstack solution template for production-ready deployments of any use case on Amazon Bedrock AgentCore.
Everything you need to know about LLM inference
A personal research and development (R&D) lab that facilitates the sharing of knowledge.
This repo contains a series of tutorials and code examples highlighting different features of the OCI Data Science and AI services, along with a release vehicle…
Deploy intelligence. Open-source infrastructure for AI agents in production.
Open-source Terraform Enterprise replacement
:rocket: Metadata tracking and UI service for Metaflow!
Local-first identity, memory, and secrets for AI agents. Portable state across models and harnesses.
Aggregates compute from spare GPU capacity
AI Infrastructure Junior Engineer Learning Track - Comprehensive curriculum for entry-level ML infrastructure engineers (0-2 years experience)
Unified AI Gateway for 30+ LLMs (OpenAI, Anthropic, Bedrock, Azure etc) with Caching, Guardrails, A/B test & cost controls. Go-native Fastest & Scalable AI Gate…
Plug-and-play homelab dashboard in one container — GPU, local-AI VRAM, Docker, systemd, host health. Built-in read-only MCP server so AI agents can explore it t…
The free, open companion to the original Grokking the System Design Interview course by DesignGurus.io.
An orchestration runtime for multi-agent AI systems. Declare agents, tools, and policies as YAML; Orloj schedules, executes, routes, and governs them for produc…
The open source, no-code MCP Server for AI-Native API Access
Efficient LLM inference on Slurm clusters.
VideoDB Python SDK
This repo is used for archiving my notes, codes and materials of cs learning.
Route inference across providers.
Context Runtime — a database query planner for LLM context. Decides what a model sees before it answers; plans it, runs it through reused substrate, and learns…
Deploy production-ready AI services in minutes. One YAML file for agents, RAG pipelines, and MCP servers — run anywhere. Inspired by docker-compose.
Persistent Claude Code agents with scheduling, sessions, memory, and Telegram.
ctx: do you remember? — a single-binary, local-first, convergent memory system for humans and machines.
Millisecond microVM sandbox forking for AI agents on Kubernetes. Firecracker VMs that restore from memory snapshots in milliseconds, fork a running VM into N co…
DevOps Projects is a curated collection of hands-on projects designed to help engineers learn and grow through real-world DevOps challenges. Inspired by platfor…
⚡ Awesome AI Gateway — curated comparison of 100+ AI gateways & LLM proxies (LiteLLM, OpenRouter, Portkey, Kong, Higress, new-api, Bifrost) by cost, security, c…
Open, self-hostable agentic runtime for organizations — durable, replayable agent sessions, human approvals, governed credentials and memory, running in managed…
OpenHermit is the open-source platform for deploying fleets of AI agents as production services — durable state, sandboxed execution, managed at scale, and the…
AI Infrastructure Performance Engineer Learning Track - GPU optimization, inference optimization, and cost reduction
Thinking in the Human · Processing in the AI · Truth in the Documentation
Self-hosted PaaS for private infrastructure. Deploy 200+ services on any server in one click. GPU monitoring, AI model management, VRAM checker, and MCP built i…
Build use cases with VideoDB
90+ production-safe OpenStack tools for AI agents via MCP. Project-scoped & read-only by default. Works with Claude, Open WebUI & any MCP client. FastMCP + Open…
Persistent memory and operational continuity for AI coding agents — append-only event ledger (AgentBook), rolling summaries, self-model, belief knowledge, and t…
Multi-model AI agent runtime. Define agents in YAML, connect 6 LLM providers, orchestrate with ReAct/Plan&Execute/Fan-Out/Pipeline/Supervisor/Swarm patterns, an…
Distributed, peer-to-peer, decentralized network for LLM inference — peers seed & leech completions, metered by BitTorrent-protocol-style upload/download ratios…
High-throughput, self-hostable LLM router in Go. BYO provider keys, k8s-native, key pooling + batch + usage observability.
NetOpsBench: Open Arena for Agentic NetOps in AI Infrastructure
Governed state engine and typed knowledge graph for AI agents, declared in YAML.
83 MCP tools for GPU infrastructure + Agent FinOps — deploy LLMs, manage VMs (Proxmox/XO/vSphere), track cost per agent, enforce budgets and model policies. Wor…
Solutions for AI Infrastructure Engineer Track
AI Infrastructure ML Platform Engineer Learning Track - Building self-service ML platforms at scale
AI Infrastructure Architect Learning Track - System design and architecture patterns for ML
