MoFlow arXiv:2609.38294 formulates agentic workflow generation as multi-objective MDP using Convex-Hull MCTS
MoFlow replaces scalar optimization of agent workflows with set-valued search that covers the Pareto front in one pass. It outperforms retrained single-objective baselines on six benchmarks while eliminating per-preference retraining. The method directly addresses the tooling limitation that forces developers to hard-code trade-offs or maintain multiple generators.
The paper introduces MoFlow to replace single-objective or weighted-sum generators that force full retraining when developer preferences shift. It models workflow construction as a multi-objective Markov decision process solved via Convex-Hull Monte Carlo Tree Search with set-valued optimistic backups. Each node maintains the full set of non-dominated outcomes, enabling post-search lookup for any linear preference vector. Evaluation reruns six single-scalar baselines for every test preference yet still records lower average hypervolume than MoFlow across mathematics, code and QA tasks.
Current agent tooling such as LangGraph and AutoGen exposes only scalar reward hooks or fixed weighting at compile time. Developers therefore embed brittle cost-accuracy compromises inside prompt templates or fine-tune separate generators per deployment tier. MoFlow removes that loop: one search run produces a lookup table that surfaces the exact workflow for a new latency budget or robustness threshold in constant time.
The approach surfaces an immediate tooling gap: existing evaluation harnesses lack native Pareto reporting, so production teams cannot audit whether deployed agents sit on the true frontier or merely on a single training slice. Integration requires only replacement of the reward aggregator inside MCTS or beam search loops already present in most workflow compilers.
Next steps include embedding MoFlow inside open agent runtimes so that preference vectors become runtime parameters rather than training constants, with hypervolume tracked as a standard CI metric.
MoFlow: Average hypervolume on math and code agent benchmarks will exceed 0.82 within nine months once set-valued search replaces scalar MCTS in at least two open-source frameworks.
Sources (2)
- [1]Primary Source(https://arxiv.org/abs/2609.38294)
- [2]Supporting Source(https://arxiv.org/abs/2303.08128)