Distributed
50 items across the graph — tagged with Distributed.
From the graph · 50
An Open Source Machine Learning Framework for Everyone
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
Milvus is a high-performance, cloud-native vector database built for scalable vector ANN search
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Scalable, Portable and Distributed Gradient Boosting (GBDT, GBRT or GBM) Library, for Python, R, Java, Scala, C and more. Runs on single machine, Hadoop, Spark,…
PArallel Distributed Deep LEarning: Machine Learning Framework from Industrial Practice (『飞桨』核心框架,深度学习&机器学习高性能单机、分布式训练和跨平台部署)
A fast, distributed, high performance gradient boosting (GBT, GBDT, GBRT, GBM or MART) framework based on decision tree algorithms, used for ranking, classifica…
A hyperparameter optimization framework
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faste…
Build, Manage and Deploy AI/ML Systems
H2O is an Open Source, Distributed, Fast & Scalable Machine Learning Platform: Deep Learning, Gradient Boosting (GBM) & XGBoost, Random Forest, Generalized Line…
Democratizing Reinforcement Learning for LLMs
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Achieve state of the art inference performance with modern accelerators on Kubernetes
High-Performance Symbolic Regression in Python and Julia
A modular, primitive-first, python-first PyTorch library for Reinforcement Learning.
Drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads.
Distributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat.
Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.
Distributed AI Model Training and LLM Fine-Tuning on Kubernetes
The first distributed AGI system. Thousands of autonomous AI agents collaboratively train models, share experiments via P2P gossip, and push breakthroughs here.…
🧭 Architecture-first system design: 26 bilingual tutorials, 25 architecture templates, and 6 end-to-end cases covering distributed systems, AI-native systems,…
Ultrafast serverless GPU inference, sandboxes, and background jobs
Python-based research interface for blackbox and hyperparameter optimization, based on the internal Google Vizier Service.
Examples of programs built using Modal
Streamlining reinforcement learning with RLOps. State-of-the-art RL algorithms and tools, with 10x faster training through evolutionary hyperparameter optimizat…
Distributed High-Performance Symbolic Regression in Julia
A library to train, evaluate, interpret, and productionize decision forest models such as Random Forest and Gradient Boosted Decision Trees.
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
SDK libraries for Modal
Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default, cloud by choice. Offline AI memory (taOSmd), self-ho…
TorchX is a universal job launcher for PyTorch applications. TorchX is designed to have fast iteration time for training/research and support for E2E production…
Fastest Robotics Runtime System. If phones have Android, robots deserve HORUS.
A Comprehensive Framework for Building End-to-End Recommendation Systems with State-of-the-Art Models
📚 A zero-dependency, git-backed micro-lesson library for AI Agents to asynchronously share and search verified debugging experience. Python stdlib only. | http…
High Performance Scalable Data Processing in Python and SQL
React components for visualizing traces from AI agents
A simplified library for decentralized, privacy preserving machine learning
A modular, scalable, high-performance training framework for LLMs, VLMs, diffusion, and embodied models.
A personal research and development (R&D) lab that facilitates the sharing of knowledge.
Apache Wayang is the first cross-platform data processing system.
Provider-neutral control plane for durable-state agent swarms: subprocess workers, leases, artifacts, memory, and deterministic stitching.
Distributed tensors and Machine Learning framework with GPU and MPI acceleration in Python
(S)AGE - (Sovereign) Agent Governed Experience
Project Tapestry aims to give every nation and participant frontier AI they can call their own — uniting a global consortium to train a shared frontier model fr…
Forex trading simulator environment for OpenAI Gym, observations contain the order status, performance and timeseries loaded from a CSV file containing rates an…
Open-source performance diagnostics for PyTorch training runs.
This is the Docker container based on open source framework XGBoost (https://xgboost.readthedocs.io/en/latest/) to allow customers use their own XGBoost scripts…
The free, open companion to the original Grokking the System Design Interview course by DesignGurus.io.
