The Question
A computer vision model on embedded hardware was slower than the team expected. Why?
The Constraint
Embedded target, existing PyTorch/C++ inference pipeline.
What We Measured
A deep-dive audit of the inference pipeline, profiled stage by stage on the target hardware.
The Finding
The real bottlenecks, which were not where the team had been optimizing.
What It Changed
Optimization effort redirected to the stages that actually dominated runtime.
