Decoding
8 items across the graph — tagged with Decoding.
From the graph · 8
Large-scale LLM inference engine
Large-scale LLM inference engine
Machine learning for NeuroImaging in Python
Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.io/aeon-7/aeon-vllm-ulti…
Code for "LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding", ACL 2024
A lean, fully-tested LLM inference server for the hardware you already have — free tiers, shared VPS, 2-core ARM boxes. OpenAI-compatible API on llama.cpp. On a…
Frequently updated list of dLLM (Diffusion Large Language Models) papers, models, and other resources
Run the AEON Bench suite on your own hardware: verified HuggingFace pull → serve → benchmark (text · agentic ×3 harnesses · vision · audio · arena · perf) → ed2…
