repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 24d ago
quic/efficient-transformers
This library empowers users to seamlessly port pretrained models and checkpoints on the HuggingFace (HF) hub (developed using HF transformers library) into inference-ready formats that run efficiently on Qualcomm Cloud AIxx (AI100, AI200 and so on) accelerators.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%FlexViT: A Flexible FPGA-based Accelerator for Edge Vision Transformers →
- PossiblePossibly related (embedding) · 49%Designing the hf CLI as an agent-optimized way to work with the Hub →
- PossiblePossibly related (embedding) · 48%Hardware startup unveils inference accelerator →
- PossiblePossibly related (embedding) · 46%Monitor and debug generative AI inference with SageMaker detailed metrics and Insights dashboard on CloudWatch →
- PossiblePossibly related (embedding) · 46%A barebones CPU-only inference engine for Qwen 3, written from scratch in pure C →
Implements
Covers
newsDesigning the hf CLI as an agent-optimized way to work with the HubnewsHardware startup unveils inference acceleratornewsMonitor and debug generative AI inference with SageMaker detailed metrics and Insights dashboard on CloudWatchnewsA barebones CPU-only inference engine for Qwen 3, written from scratch in pure C
Related across the graph
newsDesigning the hf CLI as an agent-optimized way to work with the HubpaperFlexViT: A Flexible FPGA-based Accelerator for Edge Vision TransformersnewsMonitor and debug generative AI inference with SageMaker detailed metrics and Insights dashboard on CloudWatchnewsHardware startup unveils inference acceleratornewsA barebones CPU-only inference engine for Qwen 3, written from scratch in pure C
