Mixture of agents
Mixture of agents (MoA) combines outputs from several language-model agents. In the original research method, agents operate in layers, using responses from the preceding layer as context for proposing or synthesizing a better response.
Also known as: mixture-of-agents, MoA
The original MoA paper describes multiple language-model agents producing responses, with later layers receiving earlier responses as additional context. Agents can use the same underlying model; the method does not require a different model family for every proposer.
For example, two proposers can draft answers to a policy question. A later aggregator receives both drafts and the question, then writes a combined answer. It still needs to check the policy: agreement between drafts is not independent evidence that a claim is true. Additional layers also add inference work and delay.
In episode 73 at 6:23, Wen Sang uses the MoA label while describing Genspark’s model-and-tool architecture. He discusses assigning models to tasks based on their strengths. Routing alone does not establish the paper’s layered response aggregation; the transcript does not specify Genspark’s complete implementation. This distinction matters when comparing architectures. MoA combines generated responses at the application level; mixture of experts selects expert computation within a model.
Hear it from the guest
“The models that are good at reasoning and planning, we use them to come up with the work plan for a project.”
Quotes lightly edited to remove filler words.
Sources
- Wang et al.: Mixture-of-Agents Enhances Large Language Model Capabilities — Introduces layered LLM agents that use preceding responses as auxiliary information.
Go deeper
- Together AI: MoA implementation docs
Inspect the reference code for proposing responses and combining them across layers.