Distributed Training
11 items across the graph — tagged with Distributed Training.
From the graph · 11
PArallel Distributed Deep LEarning: Machine Learning Framework from Industrial Practice (『飞桨』核心框架,深度学习&机器学习高性能单机、分布式训练和跨平台部署)
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faste…
Build, Manage and Deploy AI/ML Systems
Democratizing Reinforcement Learning for LLMs
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
TorchX is a universal job launcher for PyTorch applications. TorchX is designed to have fast iteration time for training/research and support for E2E production…
Project Tapestry aims to give every nation and participant frontier AI they can call their own — uniting a global consortium to train a shared frontier model fr…
Open-source performance diagnostics for PyTorch training runs.
This is the Docker container based on open source framework XGBoost (https://xgboost.readthedocs.io/en/latest/) to allow customers use their own XGBoost scripts…
Machine learning library, Distributed training, Deep learning, Reinforcement learning, Models, TensorFlow, PyTorch
Research platform for model training, evaluation, and experimentation across architectures, benchmarks, and recipes.
