Qwen
36 items across the graph · 1 news stories — tagged with Qwen.
Latest news
Qwen 3.8
https://www.qwencloud.com/pricing/token-plan Comments URL: https://news.ycombinator.com/item?id=48966120 Points: 514 # Comments: 368
Read full story →From the graph · 35
FastGPT is a knowledge-based platform built on the LLMs, offers a comprehensive suite of out-of-the-box capabilities such as data processing, RAG retrieval, and…
An open-source AI coding agent that lives in your terminal.
Hugging Face model with 9868 likes. Tags: transformers, safetensors, qwen3_5, image-text-to-text, conversational, license:apache-2.0, eval-results, endpoints_co…
🧑🚀 全世界最好的LLM资料总结(多模态生成、Agent、辅助编程、AI审稿、数据处理、模型训练、模型推理、o1 模型、MCP、小语言模型、视觉语言模型) | Summary of the world's best LLM resources.
Solve Visual Understanding with Reinforced VLMs
A Next-Generation Training Engine Built for Ultra-Large MoE Models
EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
A course of learning LLM inference serving on Apple Silicon for systems engineers: build a tiny vLLM + Qwen.
Hugging Face model with 3444 likes. Tags: gguf, uncensored, qwen3.6, moe, vision, multimodal, image-text-to-text, en, zh, multilingual
Hugging Face model with 2962 likes. Tags: transformers, safetensors, qwen2, text-generation, chat, conversational, en, arxiv:2309.00071, arxiv:2412.15115, base_…
Hugging Face model with 2918 likes. Tags: safetensors, qwen3_5, unsloth, qwen, qwen3.5, reasoning, chain-of-thought, Dense, image-text-to-text, conversational
TokenSpeed is a speed-of-light LLM inference engine.
Fully Open Framework for Democratized Multimodal Training
Training/Fine-tuning at the speed of light
说点啥(BiBi Keyboard):一个基于 Kotlin 的 Android 平台的 LLM 与 ASR 语音输入法键盘应用 An LLM ASR voice input method keyboard application for the Android platform based on Kotlin
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.io/aeon-7/aeon-vllm-ulti…
Run generative AI models in sophgo BM1684X/BM1688
Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime.
RL environments + evals for AI agents. Define once, train anything.
A native .NET LLM inference engine for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, and Ollama/OpenAI-compatible HTTP…
To know what models don't say out loud.
Tinybot is a lightweight personal AI Agent that is constantly evolving
Explore LLM model deployment based on AXera's AI chips
Proxy API OpenAI-compatible que usa automação com Playwright para rotear requisições para modelos do Qwen com suporte a múltiplas contas, tools e sessões persis…
A Lightweight LLM Inference Performance Simulator
A benchmark for evaluating LLM × harness performance.
A multi engine TTS & LLM edge computing playground with audio book features and more!
Run LLMs, VLMs, ASR, TTS, diarization and more fully on-device with Apple's Core AI framework (iOS/macOS 27) — one line of Swift per model, 53 models pinned to…
Chat applications that use DAG to build question-answer relationships.
Efficient multi-token attribution for reasoning language models — Python package, CLI, and HTML token traces
qMLX: Custom inference engine for Qwen 3.5 122B on Apple Silicon, extending MLX with hybrid attention support, SSD-backed KV cache, and RYS layer duplication fo…
A web-based memory usage and performance calculator for Huggingface GGUF models
Free AI models in VS Code Copilot Chat via browser auth. No API keys required.
One Mac process. Many models. Real speed. Multi-model LLM serving with prefix reuse, MTP acceleration, and OpenAI APIs — built for Apple Silicon, measured again…
