newsReddit r/MachineLearningTrust 52 · CommunityPublished 1mo agoLive · 1mo ago
Looking for feedback on a small test SLM I built completely from scratch [P]
Architecture: - Parameter count: 216.5M - Layers: 10 - Attention / no attention:** Attention — 12-head multi-head self-attention, RoPE positional encoding, SDPA. Decoder-only, pre-norm, RMSNorm + SwiGLU, tied input/output embeddings. (hidden 1032, head_dim 86, FFN 4416) - Tokenizer:** Custom 36k SentencePiece unigram, case-preserving, byte-fallback, with atomic chat/role + memory special tokens (`<|user|>`,
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference →
- PossiblePossibly related (embedding) · 48%manojmallick/sigmap →
- PossiblePossibly related (embedding) · 48%Sparse attention at million-token context →
- PossiblePossibly related (embedding) · 47%Understanding Large Language Models →
- PossiblePossibly related (embedding) · 47%NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation →
- PossiblePossibly related (embedding) · 51%Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning →
- PossiblePossibly related (embedding) · 53%On the Role of Directionality in Structural Generalization →
- PossiblePossibly related (embedding) · 52%How Much is Left? LLMs Linearly Encode Their Remaining Output Length →
Covers
paperAttend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inferencerepomanojmallick/sigmappaperSparse attention at million-token contextpaperUnderstanding Large Language ModelspaperNLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation
Covers (incoming)
paperUnlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction TuningpaperOn the Role of Directionality in Structural GeneralizationpaperHow Much is Left? LLMs Linearly Encode Their Remaining Output LengthpaperPrompt Compression via Activation AggregationpaperSuper-Tuning: From Activation-Aware Pruning to Sparse Fine-Tuningrepofkkarakurt/Nervereporiyanshibohra/TuneKit
Related across the graph
paperNLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window AdaptationpaperOn the Role of Directionality in Structural Generalizationreporiyanshibohra/TuneKitpaperSuper-Tuning: From Activation-Aware Pruning to Sparse Fine-TuningpaperUnlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction TuningpaperPrompt Compression via Activation AggregationpaperSparse attention at million-token contextpaperAttend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM InferencepaperHow Much is Left? LLMs Linearly Encode Their Remaining Output LengthpaperUnderstanding Large Language Modelsrepofkkarakurt/Nerverepomanojmallick/sigmap
