repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 6d ago
giannisanni/pulsar
SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places attention and hot experts where they fit.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 57%WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs →
- PossiblePossibly related (embedding) · 56%Alternative(s) to run CUDA on non-Nvidia hardware →
- PossiblePossibly related (embedding) · 55%We'll benchmark an Open weights LLM on any GPU you choose — drop your model + hardware and we'll run it. [D] →
- PossiblePossibly related (embedding) · 55%Top Cost-Effective Enterprise GPU Cloud Platforms for AI Workloads with H100–GB200, Elastic Scaling and Pay-as-You-Go Compute - Scott Coop →
- PossiblePossibly related (embedding) · 53%Run NVIDIA Nemotron and OpenAI GPT OSS models on Amazon Bedrock in AWS GovCloud (US) →
- PossiblePossibly related (embedding) · 64%Meet FreeToken: An Edge-Native MoE Serving Engine that Runs 753B GLM-5.2 on a Single Workstation GPU - MarkTechPost →
- PossiblePossibly related (embedding) · 52%Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2 →
- PossiblePossibly related (embedding) · 50%Accelerating Block Low-Rank Foundation Model Inference on MemoryConstrained GPUs →
Implements
Covers
newsAlternative(s) to run CUDA on non-Nvidia hardwarenewsWe'll benchmark an Open weights LLM on any GPU you choose — drop your model + hardware and we'll run it. [D]newsTop Cost-Effective Enterprise GPU Cloud Platforms for AI Workloads with H100–GB200, Elastic Scaling and Pay-as-You-Go Compute - Scott CoopnewsRun NVIDIA Nemotron and OpenAI GPT OSS models on Amazon Bedrock in AWS GovCloud (US)
Covers (incoming)
newsMeet FreeToken: An Edge-Native MoE Serving Engine that Runs 753B GLM-5.2 on a Single Workstation GPU - MarkTechPostnewsReduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2newsAccelerating Block Low-Rank Foundation Model Inference on MemoryConstrained GPUsnewsAnyone else completely tuning out these massive "open weight" drops?newsKimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost - MarkTechPostnewsPSA: DO NOT use Intel consumer platforms for multi-GPU setupsnewsOpenAI introduces ‘Ultrafast,’ a new mode that makes GPT-5.6 Sol work at 14x the speednewsQwen 3.8 2.4T at 288k tokens/s on Nvidia GB300 NVL72newsAre inference chips replacing GPUs? Investors seem to think so... [D]newsGPU Offload in Rust: Portable, Safe, and Fast
Related across the graph
newsGPU Offload in Rust: Portable, Safe, and FastnewsRun NVIDIA Nemotron and OpenAI GPT OSS models on Amazon Bedrock in AWS GovCloud (US)newsAccelerating Block Low-Rank Foundation Model Inference on MemoryConstrained GPUsnewsPSA: DO NOT use Intel consumer platforms for multi-GPU setupsnewsAnyone else completely tuning out these massive "open weight" drops?newsKimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost - MarkTechPostnewsQwen 3.8 2.4T at 288k tokens/s on Nvidia GB300 NVL72paperWattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMsnewsReduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2newsWe'll benchmark an Open weights LLM on any GPU you choose — drop your model + hardware and we'll run it. [D]newsMeet FreeToken: An Edge-Native MoE Serving Engine that Runs 753B GLM-5.2 on a Single Workstation GPU - MarkTechPostnewsAre inference chips replacing GPUs? Investors seem to think so... [D]newsOpenAI introduces ‘Ultrafast,’ a new mode that makes GPT-5.6 Sol work at 14x the speednewsAlternative(s) to run CUDA on non-Nvidia hardwarenewsTop Cost-Effective Enterprise GPU Cloud Platforms for AI Workloads with H100–GB200, Elastic Scaling and Pay-as-You-Go Compute - Scott Coop
