Glossary
Inference
Running a trained model to produce output — the operational, per-request phase of AI, as opposed to training. Latency, throughput and cost per token are its economics.
In AGNT
AGNT points inference anywhere: frontier APIs, fast-inference clouds, or fully local through Ollama.