research

Humanize: Multi-Agent System Boosts Reliability of AI-Generated Code

Summarized by AI from reporting by ArXiv cs.AI, published under our editorial policy.

Researchers introduced Humanize, a multi-agent orchestration workflow that improves the reliability of AI-generated code by enforcing explicit, mechanically enforced decisions at the boundaries between planning, implementation, review, and learning.

A flowchart illustrating the Humanize workflow, showing the stages of planning, implementation, and review.

Key takeaways

  • Humanize is a multi-agent orchestration workflow designed to improve the reliability of agentic coding.
  • It uses judgement engineering to enforce explicit decisions at key stages: planning, implementation, review, and learning.
  • Humanize involves a human approving a plan contract, a builder agent implementing it, and a reviewer agent from another vendor deciding completion.

Researchers from various institutions released Humanize, a new multi-agent orchestration workflow designed to improve the reliability of agentic coding. Agentic coding, where AI agents generate code, is cost-effective but often produces unreliable results because the same agent that writes the code also judges its completeness.

Humanize addresses this issue by introducing judgement engineering, a method that enforces explicit, mechanically enforced decisions at key stages: planning, implementation, review, and learning. This structured approach ensures that different agents handle different tasks, reducing the risk of errors.

Humanize's Three-Role Workflow: Planner, Builder, Reviewer

Humanize's workflow involves three main roles: a planner, a builder, and a reviewer. A human first approves a plan contract, which outlines the requirements and constraints for the coding task. A builder agent then implements the plan in rounds, making incremental progress. Finally, a reviewer agent from a different vendor assesses the completion of the task, ensuring objectivity and reliability. Deterministic hooks, rather than a model, manage the transitions between these stages, providing a clear and consistent process.

Why Judgement Engineering Improves Code Quality

The key innovation of Humanize is its focus on judgement engineering. By separating the planning, implementation, and review stages, Humanize ensures that each step is handled by a specialized agent or human, reducing the likelihood of errors. This structured approach is particularly important in coding, where small mistakes can lead to significant issues. Humanize's use of deterministic hooks also ensures that the transitions between stages are clear and consistent, further enhancing reliability.

Implications for Developers and Users

For developers, Humanize offers a more reliable way to generate code using AI agents. By ensuring that each stage of the coding process is handled by a specialized agent or human, Humanize reduces the risk of errors and improves the overall quality of the output. This can be particularly useful for complex coding tasks that require careful planning and review. For users, Humanize's structured approach can provide greater confidence in the reliability of AI-generated code, making it a valuable tool for both professional developers and hobbyists.

Frequently asked

What is agentic coding?
Agentic coding is a process where AI agents generate code. It is cost-effective but often produces unreliable results because the same agent that writes the code also judges its completeness.
How does Humanize improve reliability?
Humanize improves reliability by separating the planning, implementation, and review stages, ensuring that each step is handled by a specialized agent or human.
Is Humanize available for public use?
As of now, Humanize is still in the research phase and not widely available for public use.