Closed beta · early access

Multi Agent Debatesby Delibora

A structured multi-agent debate platform for decisions, pitches, research, and ideas. Give several AI participants the same question, compare their arguments, and leave with an inspectable result.

Request free beta accessExplore the researchFree during beta · no credit card
Participants
2 – 100
Formats
10
Model routes
Cloud + local
01 / Scientific foundation

Protocols connected to their research origins.

The platform's debate workflows are informed by published multi-agent research. The references below explain the methods and tradeoffs behind structured debate, judging, and collaboration on arXiv, ACL Anthology, or NeurIPS.

New to the field? Read our guide on when to use multi-agent debate vs self-consistency.

  1. Improving Factuality and Reasoning in Language Models through Multiagent Debate

    Du, Li, Torralba, Tenenbaum, Mordatch

    Foundational result: agents critiquing each other across rounds converge on more factual, better-reasoned answers.

  2. Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

    Liang, He, Jiao, Wang, Wang, Wang, Yang, Shi, Tu

    Introduces the Degeneration-of-Thought problem and shows multi-agent debate unlocks divergent thinking where self-reflection stalls.

  3. Paper 03ACL 2025
    M-MAD: Multidimensional Multi-Agent Debate for Translation Evaluation

    Feng, Zhao, Lyu, Li, Tu, Wang

    Introduces the per-dimension arbiter sweep that informs the platform's truth-seeking verdict scoring.

  4. Paper 04ICLR 2024
    ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

    Chan, Chen, Yu, Lu, Sun, Liu

    Demonstrates that multi-agent debate panels evaluate generated text more reliably than single-judge baselines.

  5. Debating with More Persuasive LLMs Leads to More Truthful Answers

    Khan, Hughes, Valentine, Ruis, Sachan, Radhakrishnan, Bowman, Perez

    Strong empirical evidence that debate makes weaker judges reliably select truthful answers from stronger debaters.

  6. Paper 06NeurIPS 2025
    Multi-Agent Debate for LLM Judges with Adaptive Stability Detection

    Hu, Tan, Wang, Qu, Chen

    Formalizes debate among LLM judges and adds adaptive stability detection so debates stop when consensus stabilizes — improving accuracy over majority vote.

  7. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

    Wu, Bansal, Zhang, Wu, Li, Zhu, Wang, Saied, Awadallah, Awadalla, Wang

    Shows that role-specialized agent groups orchestrated through structured conversation consistently outperform monolithic prompts on complex tasks.

  8. Reflexion: Language Agents with Verbal Reinforcement Learning

    Shinn, Cassano, Berman, Gopinath, Narasimhan, Yao

    Studies verbal self-critique loops in which an agent reflects on prior attempts, generates feedback, and retries.

  9. Self-Refine: Iterative Refinement with Self-Feedback

    Madaan, Tandon, Gupta, Hallinan, Gao, Wiegreffe, Alon, et al.

    Single-model iterative refinement via self-generated feedback. The minimal version of what multi-agent debate scales up across roles.

  10. Mixture-of-Agents Enhances Large Language Model Capabilities

    Wang, Bai, Liu, Chen, Cardie, Zhang, et al.

    Layered multi-LLM collaboration where each layer's agents refine the previous layer's outputs. Open-source MoA reaches 65.1% on AlpacaEval 2.0, beating GPT-4 Omni.

  11. Paper 112026
    Demystifying Multi-Agent Debate: The Role of Confidence and Diversity

    Choi, Zhu, Li, et al.

    Pinpoints when multi-agent debate actually beats majority vote: diversity-aware initialization plus calibrated confidence communication.

Further reading

02 / Capabilities

Ten formats for ten different jobs.

Start with a purpose-built workflow for decisions, expert review, debate, dialogue, team competition, pitches, qualitative research, audience simulation, idea selection, or a custom setup.

  • Decision stress-test

    2 opposing agents

    Two opposing agents pressure-test the assumptions and converge on a practical recommendation.

    Output: Decision memo

  • Expert panel

    3–5 experts

    A small panel debates your topic and distills the most useful insights.

    Output: Key insights, no verdict

  • Judged debate

    2 debaters + arbiter

    Two debaters build opposing cases while an arbiter scores the exchange.

    Output: Arbiter's scored verdict

  • 1:1 Human Dialogue

    2 Worker-backed LLMs

    Two Worker-backed LLMs take turns in a private chat, each believing the other speaker is human.

    Output: Dialogue transcript

  • Team Battle

    2 saved Teams + fixed Arbiter

    Every member prepares privately, saved spokespersons debate in a fixed public sequence, and a neutral Arbiter scores the result.

    Output: Five-dimension scorecard and verdict

  • Shark Tank

    2–4 Workers

    Choose what is being decided, then have 2–4 Workers judge the pitch independently against that goal.

    Output: Cited scorecard and verdict

  • Focus group

    4–12 Workers

    A moderator follows an editable guide with known Worker personas and grounded findings.

    Output: Themes and verified quotes

  • TribeMind

    12–50 synthetic participants

    Generated participants react independently, then influence each other in networked rounds.

    Output: Metrics and Observer report

  • Idea Tournament

    4–16 ideas, 1–5 judges

    Ideas are seeded, then argued head to head in a bracket until one survives with the best of the rest folded in.

    Output: Champion spec and kill cards

  • Custom discussion

    2–12 participants

    Choose the protocol, participants, models, runtime, evidence, and notifications.

    Output: Configurable

  • 01

    Evidence packs & internet research

    Attach source material to discussions. Truth-Seeking Debate can also prepare dedicated research for both sides and the neutral Arbiter.

  • 02

    2 to 100 reasoning agents

    Start with a focused pair or assemble a larger panel. Specialized formats add their own participant limits and roles.

  • 03

    Workers, Teams, Personas & Playbooks

    Reuse configured participants, saved teams, visible role presets, shared rules, and hidden turn guidance across discussions.

  • 04

    Format-specific outputs

    Get a decision memo, scored verdict, qualitative report, audience metrics, dialogue transcript, or champion spec—depending on the format.

  • 05

    Lab Experiments

    Copy a draft into hidden child runs, sweep sampling parameters, score each transcript, and stop on score, iteration, or cost limits.

  • 06

    Live human intervention

    Inject a human message into a running or paused discussion when the agents need a correction, constraint, or new piece of context.

  • 07

    OpenRouter + LM Studio

    Use cloud models through OpenRouter or local models through LM Studio, with configurable model fallbacks.

  • 08

    Runtime controls

    Bound work with turn, wall-clock runtime, and observed-cost limits so long-running discussions stop predictably.

  • 09

    Durable execution

    Dedicated workers advance discussions, pitches, focus groups, simulations, evidence, and research outside long HTTP requests.

  • 10

    Transcripts, logs & exports

    Review persisted turns, run logs, token and cost metadata, then export portable JSON records for further analysis.

  • 11

    In-app + email notifications

    Receive an in-app notification—and optional Resend email—when a run finishes or reaches its cost cap.

03 / Pitch and research

Go beyond a debate transcript.

Multi Agent Debates includes dedicated workflows for pitch evaluation, moderated qualitative research, synthetic-audience simulation, and head-to-head idea selection.

  • Shark Tank2–4 Workers

    Get a scored verdict on a pitch for a named audience

    Choose what is being decided, then have 2–4 Workers judge the pitch independently against that goal.

    Output

    Cited scorecard and verdict

  • Focus group4–12 Workers

    Run moderated qualitative research

    A moderator follows an editable guide with known Worker personas and grounded findings.

    Output

    Themes and verified quotes

  • TribeMind12–50 synthetic participants

    Simulate broader audience reactions and opinion shifts

    Generated participants react independently, then influence each other in networked rounds.

    Output

    Metrics and Observer report

  • Idea Tournament4–16 ideas, 1–5 judges

    Choose one idea out of many and know why

    Ideas are seeded, then argued head to head in a bracket until one survives with the best of the rest folded in.

    Output

    Champion spec and kill cards

04 / Evidence and iteration

Test settings. Ground the result.

Lab Experiments compare parameter choices across child runs, while Evidence Packs and optional Truth-Seeking internet research keep factual work connected to source material.

Lab Experimentsparameter sweeps

Compare configurations systematically.

Start from a draft discussion, create hidden child runs, vary temperature and repetition, frequency, and presence penalties, then score each transcript against a validation prompt and expected outcome.

  • Inspect child-run links, scores, costs, and parameter snapshots
  • Pause, resume, stop, edit, or delete an Experiment
  • Stop on score threshold, iteration limit, or total cost cap
Evidence and researchsource-aware

Give agents material they can inspect.

Upload an Evidence Pack for a discussion. In Truth-Seeking Debate, optional internet research prepares separate material for both debaters and neutral material for the Arbiter before the relevant phases run.

  • Shared files and pasted evidence stay attached to the draft
  • Research is purpose-built for Truth-Seeking Debate
  • Claims and citations remain visible in the persisted run
What this adds

Compare runs without losing the transcript, parameters, evidence, or cost trail behind the result.

Join the beta
05 / Outputs

Every format leaves something inspectable.

Multi Agent Debates keeps the transcript and operational trail, then adds the output that fits the job: a memo, verdict, report, scorecard, dialogue, metrics package, or tournament result.

  • Decisions and verdicts

    Know what was decided—and why.

    Decision stress-tests produce a decision memo. Judged Debate returns an Arbiter score, and Team Battle ends with a five-dimension scorecard and verdict.

  • Qualitative research

    Turn the session into findings.

    Focus Group produces grounded themes and verified quotes. Expert Panel distills key insights without forcing a winner.

  • Audience and pitch signals

    Inspect reactions, scores, and movement.

    Shark Tank returns a cited pitch scorecard and verdict. TribeMind provides descriptive metrics and a grounded Observer report.

  • Selection and exploration

    Keep the path to the result.

    Idea Tournament produces a champion spec and kill cards. Human Dialogue preserves the private-chat transcript, while Custom Discussion keeps its output configurable.

  • Persisted transcripts and run logs
  • Token and cost metadata
  • Format-specific reports and scorecards
  • Portable JSON exports
06 / Use cases

A workspace for questions that benefit from structured disagreement.

Multi Agent Debates can support message testing, qualitative research, pitch review, argument mapping, product decisions, audience simulation, and idea selection.

Political Campaigns

Test a message from more than one angle.

Use Team Battle for structured opposition, Focus Group for moderated reactions, or TribeMind to observe synthetic audience shifts across rounds.

Academic Research

Run a simulated peer review.

Put a hypothesis in front of an Expert Panel or compare two competing interpretations in a Judged Debate with attached evidence.

Marketing & Brand

Move from reactions to a usable report.

Moderate known personas in a Focus Group, score a campaign pitch in Shark Tank, or simulate wider audience response with TribeMind.

Legal & Due Diligence

Lay out competing arguments.

Use a Decision Stress-Test or Judged Debate to compare competing cases, preserve the transcript, and keep supporting evidence with the run.

Product Strategy

Stress-test a decision or select an idea.

Challenge a roadmap choice with opposing agents, or send several concepts through an Idea Tournament for a champion spec and kill cards.

Education & Training

Make contrasting viewpoints visible.

Review the transcript from an Expert Panel, 1:1 Human Dialogue, or Judged Debate and discuss how each participant handled the question.

Investigative Journalism

Rehearse the counterargument before publication.

Attach source material, run a Judged Debate or Team Battle, and collect the counterarguments your draft still needs to answer.

Just for Fun

Stage a conversation or bracket.

Put two AI personalities into a 1:1 Human Dialogue, build a custom discussion, or let several ideas compete in a tournament.

Frequently asked

Multi-agent debate, demystified.

Quick answers to the questions teams ask before they run their first Multi Agent Debates session.

01What is multi-agent debate?

Multi-agent debate is an AI reasoning technique where multiple language models, or multiple instances of the same model with different roles, discuss a question across structured rounds. Research suggests it can improve some reasoning and evaluation tasks, though results depend on the models, format, and problem.

02How is Multi Agent Debates different from prompting ChatGPT or Claude?

A single prompt gives you one model's first-pass answer. Multi Agent Debates gives you 10 purpose-built formats: Decision stress-test, Expert panel, Judged debate, 1:1 Human Dialogue, Team Battle, Shark Tank, Focus group, TribeMind, Idea Tournament, and Custom discussion. Each format preserves the run and produces an output suited to the job.

03Which AI models can I use with Multi Agent Debates?

Multi Agent Debates supports models available through OpenRouter and local models served by LM Studio. You can select models per Worker and configure automatic fallback order in Settings.

04Is multi-agent debate scientifically validated?

The platform's debate workflows draw on published research from several institutions. The evidence is promising for some reasoning and evaluation tasks, but results remain model-, protocol-, and task-dependent. Key papers include:

See all 11 foundational papers ↓

05Where can I read more about multi-agent debate?

We publish free, in-depth guides on multi-agent debate methodology—no signup required. Start here:

See all 11 foundational papers ↓

06What can I use Multi Agent Debates for?

Use it to pressure-test decisions, gather specialist viewpoints, compare two arguments, observe a private AI-to-AI dialogue, stage Team Battles, evaluate pitches, run moderated Focus Groups, simulate audience response with TribeMind, select ideas in a tournament, or configure a discussion from scratch.

07How does the Focus group format work?

Choose 4–12 Workers with known personas, edit the discussion guide, and let a neutral moderator run the session. The resulting report is grounded in that simulated panel and includes themes and verified quotes rather than presenting the run as real population research.

08What do Shark Tank, TribeMind, and Idea Tournament produce?

Shark Tank returns a cited pitch scorecard and verdict. TribeMind runs 12–50 fictional participants through independent and social rounds, then produces descriptive metrics and a grounded Observer report. Idea Tournament compares 4–16 ideas head to head and produces a champion spec plus kill cards.

09What do Lab Experiments do?

Lab Experiments copy a draft discussion into hidden child runs, vary temperature and repetition, frequency, and presence penalties, then evaluate each transcript against your validation prompt and expected outcome. They can stop on a score threshold, iteration limit, or total cost cap.

10Is Multi Agent Debates an alternative to AutoGen, CrewAI, or LangGraph?

Multi Agent Debates focuses on purpose-built discussion formats, durable run monitoring, and format-specific reports rather than general-purpose agent orchestration. AutoGen, CrewAI, and LangGraph are broader frameworks for building custom agent applications and graphs.

11How much does a run cost?

Cost varies with the selected models, participant count, number of rounds, research calls, and report generation. Multi Agent Debates records token and cost metadata and supports turn, runtime, and observed-cost limits. Lab Experiments also support their own total cost cap.

12Can I use Multi Agent Debates without writing any code?

Yes. All 10 discussion formats, Workers, Teams, Personas, Playbooks, Lab Experiments, live monitors, and exports are available through the web interface.

13Is my data used to train AI models?

Multi Agent Debates does not train models on your transcripts. Prompts are sent to the model route you configure, so provider data handling depends on your OpenRouter model choice or LM Studio setup. The platform persists workspace records in its configured Supabase storage and provides archive, delete, and export controls.

14What is the Degeneration of Thought problem?

Degeneration of Thought, formalized by Liang et al. (EMNLP 2024), is a failure mode where an LLM commits to an answer and struggles to produce genuinely novel reasoning during self-reflection. Separating roles across agents may help surface different arguments, but it does not eliminate the underlying risk.

07 / Beta access

See whether Multi Agent Debates fits your way of working.

Closed beta is rolling out gradually. Tell us what you would use it for, and we will send an invite when there is room.

Step 1 of 3

Your field

  1. Step 1, current step
  2. Step 2
  3. Step 3
What industry or niche are you in?

Choose the field that best describes your work.