newsReddit r/MachineLearningTrust 52 · CommunityPublished 1mo agoLive · 1mo ago
PyTorch model running 170x slower on T4 vs A100. What could cause a bottleneck this extreme? [D]
Hey everyone, Seeing a ~170× slowdown running a point-tracking model on an NVIDIA T4 compared to an A100. On A100 the tracker takes ~0.5 seconds per half-video. On T4 the same call takes ~85 seconds. Video is 47 frames at 256×256, batch 1. I expect a meaningful gap between these cards, but 170× feels too large to explain by generational hardware differences alone. Setup: Precision: pure FP32 Architecture: builds local 4D corre
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 46%GPU Parallelization Strategies for Forward and Backward Propagation in Shallow Neural Networks: A CUDA-Based Comparative Study →
