Speculative Decoding
6 items across the graph — tagged with Speculative Decoding.
From the graph · 6
repo
dphnAI/sonar
→repoLarge-scale LLM inference engine
dphnAI/aphrodite-engine
→repoLarge-scale LLM inference engine
AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-DFlash
→repoFully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.io/aeon-7/aeon-vllm-ulti…
facebookresearch/LayerSkip
→repoCode for "LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding", ACL 2024
swellweb/reame
→repoA lean, fully-tested LLM inference server for the hardware you already have — free tiers, shared VPS, 2-core ARM boxes. OpenAI-compatible API on llama.cpp. On a…
AEON-7/Aeon-Bench-Pod
→Run the AEON Bench suite on your own hardware: verified HuggingFace pull → serve → benchmark (text · agentic ×3 harnesses · vision · audio · arena · perf) → ed2…
