Definition
Training teaches a model; inference applies it. When ChatGPT answers your question or a classifier labels an email, that is inference.
Production inference concerns latency, cost, scaling, and reliability — often separate from the research notebook where the model was built.
In simple terms
Training is studying for an exam. Inference is taking the exam — applying what you learned to new questions.
Where you see it
- API calls to OpenAI are inference.
- On-device keyboard suggestions run inference locally.
- Somali ASR transcribing audio in real time.
How it works
1.Load checkpoint
Serve trained weights in a runtime (PyTorch, ONNX, vLLM).
2.Accept input
Batch requests for efficiency.
3.Forward pass only
No backpropagation — faster than training.
4.Return output
Text, labels, embeddings, or audio depending on the model.
Why it matters
- Most users interact with AI at inference time — speed and cost here define product quality.
Often confused
Bigger models always infer faster.
Larger models are usually slower; distillation and quantization trade size for speed.