AI Glossary

LLMOps

LLMOps is the practice of deploying, evaluating and maintaining large-language-model applications. It covers the model and surrounding components, including prompts, retrieval, tools, quality monitoring and operating cost.

· Updated · Chain of Thought

A support assistant can change behavior when its prompt changes, its model route changes or new documents enter its retrieval index. A team records those versions, runs representative evaluations before a release, and uses traces to investigate failures after deployment.

For example, a new prompt may make replies shorter while causing the assistant to omit an important returns-policy exception. A latency monitor will not detect that quality regression. An evaluation with policy exceptions can, while a trace shows which passages and tool results the model received.

MLOps supplies related practices for model development and operation. LLMOps extends the operational view to components a team may control even when it does not train the underlying model. This matters for recovery: reverting a prompt alone may not restore behavior if the retrieval index or tool contract also changed. Plan and test releases as application changes, with task outcomes as well as cost and speed visible.

Sources

  • MLflow: Agents and LLMs — Documents tracing, evaluation, prompt management and deployment across language-model application development.

Go deeper