LLM Council: how to ask multiple AIs at once and get one answer you can use
Last updated: 6 October 2026. By Leonardo Cantoni, founder of Rileva.
An LLM council is a way of asking several AI models the same question independently, then comparing their answers to produce one final answer that shows where the models agree, where they disagree, and why. Each model answers on its own, without seeing the others. A final step reads all the answers side by side and writes the conclusion. The point is not to count votes. It's to surface the disagreement a single model would hide.
The term took off after Andrej Karpathy published a small open-source project called llm-council, which sends one query to several models, has them review each other anonymously, and has a "Chairman" model write the final response. The idea itself is older and simpler: get a second opinion, and then a third.
You can do it for free by hand in a few browser tabs (method below), or use a tool that does it for you.
How an LLM council works
Most LLM councils, including hand-made ones, follow three steps:
- Independent answers. The same question goes to two or more models, each in a fresh context. None of them sees what the others wrote. That independence is what makes the exercise worth doing.
- Comparison. The answers are compared: which ones reach the same conclusion, which ones don't, and what assumption or priority explains the difference. Good setups do this blind, hiding which model wrote which answer so brand reputation doesn't tip the scales.
- One final answer. A single answer is written from the comparison. A useful one doesn't just average the inputs. It gives a direct conclusion, states the shared ground once, names the real disagreements, and says what the right choice depends on.
When an LLM council helps, and when it doesn't
An honest summary: a council is a tool for questions where judgment matters and a confident wrong answer would cost you something. For many everyday questions it's overkill.
It tends to help when:
- The question involves judgment, not lookup. Strategy, pricing, a hire, choosing between two plans, reviewing an argument. There's no single correct answer, and different framings lead to different recommendations.
- The stakes are high enough to want a second opinion. If you'd ask a colleague to sanity-check it, asking two more models is a cheap version of that.
- You suspect one model is just agreeing with you. A model that has spent a whole chat helping you build a case is anchored to it. A fresh model with only the brief isn't.
- You want the risks someone else would spot. Often one answer raises a risk the others miss. Seeing it is worth more than any amount of agreement.
It usually doesn't help (or isn't worth it) when:
- The answer is a fact you can check. For a date, a definition or a formula, look it up or verify it directly. Three models agreeing on a fact doesn't make it true.
- The models share the same blind spot. Models are trained on overlapping data and can be wrong in the same way, especially about recent events or niche topics. Agreement means "consistent", not "correct". For anything current, you need sources, not more opinions.
- You need one voice. For a piece of creative writing or an email in your tone, blending three styles usually makes it worse. Pick one model and iterate.
- Speed or cost matters more than confidence. A council takes longer and uses more compute than a single answer. For quick, low-stakes questions, one good model is fine.
A council won't make AI infallible. What it changes is that the uncertainty becomes visible instead of hidden behind one fluent answer.
How to run an LLM council by hand (free method)
You need accounts with two or three different AI chat apps (free tiers are fine for most questions) and about ten minutes.
Step 1: write one brief
Write your question once, with the context a smart outsider would need: what you're deciding, your options, your constraints, and what success looks like. Keep it in a text file so every model gets the same version.
End the brief with an instruction that forces each model to commit:
Give me a clear recommendation. Then list the assumptions it rests on,
and the specific conditions under which your recommendation would be wrong.
Don't hedge to appear balanced, and don't give generic advice.
That last part matters. Without it, models tend to give balanced overviews that are hard to compare.
Step 2: ask three models in fresh chats
Paste the same brief into three different models, each in a new chat with no earlier history. Don't show any of them the others' answers, and don't paste in your own earlier reasoning. You want independent opinions, not three versions of yours.
Models from different companies are a better choice than three models from the same one. They're more likely to be trained and tuned differently, though there's no guarantee they'll disagree.
Step 3: compare blind
Open a fourth chat. Paste the three answers, labelled only A, B and C (remove anything that reveals which model wrote which), and use this prompt:
Here are three independent answers to the same question, labelled A, B and C.
Compare them and write:
1. The short answer: your own recommendation in one or two sentences.
2. Where they agree: the shared substance, stated once.
3. Where they differ: what each one concluded and why (the assumption,
priority or risk weighting behind it).
4. What it depends on: the facts about my situation that decide which view is right.
Don't pick a position because more answers hold it; a minority view can carry
the strongest argument. Flag any important risk only one of them raised.
Read the "where they differ" section first. That's where the useful information usually is.
Tip: the comparing model may also have been one of the three that answered. Labelling the answers A/B/C at least means it doesn't know which one is its own.
If you code: self-hosted options
Karpathy's llm-council repo is a small local web app that does the full loop through OpenRouter (you need an OpenRouter API key). Its README describes it as a weekend "vibe coded" project that he doesn't plan to support, so treat it as a starting point to modify, not a product.
How Rileva's Council does it
The by-hand method works. It's also four tabs and a lot of pasting, which is why most people skip it. We built Council in Rileva to do the same loop from one message. This is how it works, as described in the product:
- You send the question once. In the composer, switch from Auto to Council and send.
- Three models answer independently. Each one gets your question (and the earlier conversation as context) and is never told that other models exist. Each is told to state a clear conclusion, the assumptions it rests on, and when it would be wrong, and not to hedge just to look balanced. Rileva tries to seat models from different providers before it repeats one. With Custom Council you can choose which providers sit on it.
- The comparison is blind. The answers are passed on as "Perspective A, B, C". The comparison step groups them by conclusion and writes the final answer without knowing which model wrote which. Real model names are put back only after it's finished.
- You get one answer in a fixed shape: the short answer, where they agree, where they differ (and why), and what it depends on.
- It is explicitly not a vote. The comparison step is told never to pick a position because more perspectives hold it, never to invent consensus, and to call out a risk that only one perspective raised.
- Every perspective stays readable. Open the drawer to read each model's full answer in its own words.
To be honest about the limits: the comparison step is also an AI model, and it may come from the same family as one of the members. It just doesn't know which answer is which. Council also takes longer than a single answer. It doesn't make answers correct, and it's not a fact-checker. For current facts, Rileva's web research answers come with sources you can open.
For everything that doesn't need a council, leave the composer on Auto: Rileva picks a model for each request based on the kind of work, shows which model answered, and lets you lock a model whenever you want.
Other ways to ask multiple AIs at once
You have options beyond Rileva, and some may suit you better:
- Side-by-side chat apps (for example ChatHub, Poe's multi-bot chat, or TypingMind's multi-model chats) send one prompt to several models and show the answers next to each other. That's great when you want to read and judge them yourself.
- Perplexity's Model Council runs a query across three models and has a synthesizer show where they agree and differ. Perplexity's February 2026 launch post said it was available to Max subscribers on the web, so check its current plans.
- By hand, as above. Free, slower, and fully under your control.
What sets a council apart from a plain side-by-side view is the last step: someone (you, or a model working blind) does the comparing and commits to a conclusion.
FAQ
What is an LLM council? An LLM council is a method where several AI models answer the same question independently, and their answers are then compared to produce one final answer showing where the models agree, where they differ, and why. It's a structured way to get a second and third opinion from AI.
Who came up with the term "LLM council"? The term became popular through Andrej Karpathy's open-source llm-council project on GitHub. It sends a query to several models, has them review each other's answers anonymously, and has a "Chairman" model write the final response. The underlying idea, independent opinions followed by a comparison, is much older.
Is an LLM council more accurate than a single AI? Not automatically. A council makes disagreement and uncertainty visible, which helps you catch weak reasoning and missed risks. But models can share the same blind spots, so agreement between them doesn't prove an answer is correct. For facts, check sources.
How do I ask multiple AIs at once for free? Write one brief, paste it into two or three different AI chat apps in fresh chats, then paste their answers, labelled A, B and C, into a fourth chat and ask it to compare: where they agree, where they differ and why, and what the answer depends on. Free tiers are usually enough.
Is an LLM council the same as majority voting? No. A good council doesn't pick the answer most models gave. A minority answer can carry the strongest argument or the most important risk. The value is in explaining why the models differ, not in counting them.
When should I not use an LLM council? Skip it for simple factual lookups, for current events (use sources instead), for creative work that needs one consistent voice, and for quick low-stakes questions where speed matters more than a second opinion.
How is Rileva's Council different from a side-by-side chat app? Side-by-side apps show you several answers and leave the comparing to you. Rileva's Council asks three models independently, compares their answers blind (without knowing which model wrote which), and writes one answer laid out as the short answer, where they agree, where they differ, and what it depends on. Every full answer stays available to read.
Can I choose which models are on the council? In Rileva, the default Council picks three models for the task and tries to use different providers first. Custom Council lets you choose the providers yourself. In the by-hand method you choose everything.
Try it
The easiest way to see whether a council is worth it for you is to run one on a real question you're stuck on. Use the by-hand method above, or try it in Rileva free (no credit card required). The link below opens a chat with an example prompt already in the composer. Replace the brackets, switch the composer to Council, and send.
Try Council on your own question →
Prefilled prompt (decoded): "I'm deciding between [option A] and [option B]. My situation: [2-3 lines of context]. Give me a clear recommendation, the assumptions it rests on, and the specific conditions under which it would be wrong."
If you try it, I'd like to hear where the models disagreed. That's usually the most interesting part.
Related: Poe alternative · Pricing

