newsReddit r/MachineLearningTrust 52 · CommunityPublished 28d agoLive · 27d ago
Tried testing qwen 35b moe model on s26 ultra , without compromising on precision [R] ,[D]
Started testing a private qwen 35B moe capacity LLM runtime on s26 ultra, early testing shows that active model footprint can fit within the device’s memory limits.( not sharing the methods or architecture used) and results suggest roughly 90 input processing t/s achievable after optimisation and output generation is around 8 tokens/s on this mobile. Point is i learned ai ml based on my interest and no formal PhD , I have compute and resources to test. An
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%luziyao1995/vllm →
- PossiblePossibly related (embedding) · 53%abhisadineni/vllm →
- PossiblePossibly related (embedding) · 52%alibaba/rtp-llm →
- PossiblePossibly related (embedding) · 52%vllm-project/vllm →
- PossiblePossibly related (embedding) · 51%Evaluate a model properly →
- PossiblePossibly related (embedding) · 51%Helldez/BigMoeOnEdge →
