Guide
What Actually Makes a Good AI Twin (And What Doesn't)
The usual first attempt
You upload your documents to a general AI tool and tell it to respond as you. The result is accurate enough and sounds nothing like you. You adjust the prompt a few times, it improves slightly, and after a fortnight you stop opening it.
The idea is sound. What goes wrong is that most implementations model the wrong thing. They capture what a person has written. A useful twin has to capture how that person decides.
What bad twins look like
They go generic off-script. Inside the material they were given, they are fine. One step outside it and you get the balanced, polite, "there are several factors to consider" answer every general model produces.
They drift. Ask on Monday whether to take on a client who wants a fixed fee for open-ended work and the answer is no. Ask on Thursday and it is "it depends". Nothing stable sits behind the replies.
They know about you, not how you decide. They can quote your article on pricing. They cannot say what you would do with a client who accepts the price and then contests every invoice.
They cannot explain themselves. Ask why and you get the same answer in different words. There is no reasoning to show, because the reply was the most plausible text, not a conclusion drawn from anything.
They invent your opinion. Perhaps the worst one. Asked about something you have never addressed, a bad twin supplies a confident view in your name.
What good twins do
1. They reason from decisions, not descriptions
"Values long-term relationships" describes most advisers and predicts nothing. "Absorbs a small overrun once for a client in their first quarter, and raises it at the next review; never absorbs a second" is something a twin can apply.
So the first mark of a good twin is what it was built from. On Imora that is the reasoning graph: your answers to real scenarios, your corrections and your documents are kept as evidence, and the evidence is distilled into principles with their conditions and exceptions. Introducing the reasoning graph goes through the mechanics.
2. They show their work
A reader should be able to see why the twin said what it said. Imora's replies open with the twin's read of the situation, and "why this answer" lists the principles and evidence used and how confident it could be. If a document of yours was used, it is cited.
This is not decoration. It is how your team learns to trust some answers and check others, and it is how you spot a principle that has been learned wrongly.
3. They admit gaps
When a shared Imora twin has no recorded view on a question, it says so. It can offer an approach, clearly framed as an approach and not your position, and the question lands in your inbox of questions to answer.
Some people see this as the twin failing. We see the opposite. Your name is on the answer. A twin that says "she has not recorded a view on this" protects you. One that improvises does not.
4. They keep corrections
If you correct a twin and it makes the same mistake next week, it is a toy. On Imora, marking a reply "Not what I'd say" and giving your version creates evidence that outranks everything else. "Sounds like me" confirms a good reply.
Equally important is who can do this. Only the owner teaches the twin. Nothing a visitor says changes how it thinks.
5. They adapt to the audience without changing the position
Your view on firing a client should be the same whether a junior or the client's peer is asking. How you put it will differ. A good system separates the two: one model of how you decide, and twins that set audience, focus and tone. Build once, deploy many covers that design.
Measuring it: completeness and accuracy are different things
"Does it sound like me?" is a poor test, especially when you are the one judging. There are two separate questions and they need separate measures.
How much does it have to go on? On Imora that is Mirror Score. It rises as you add scenarios, corrections, interview sections and documents. It is a completeness measure. It is not an accuracy score, and a high Mirror Score does not prove the twin is right.
Does it get you right? That is the fidelity check. We hide some of your real answers from your twin, then check which version gets closest to what you actually said. Because the twin never saw those answers, it cannot simply repeat them.
Whatever tool you use, ask which of these a headline score is measuring. Most measure the first and imply the second.
Why calibration matters
You are not the same thinker you were six months ago. You lost a pitch and changed how you qualify leads. You made a bad hire and now weigh references differently.
A twin that is never updated does not degrade. You move and it stays put. Short calibration sessions, about five minutes each, put new scenarios in front of you, aimed at the thin parts of your graph. Your answers become evidence and the principles are re-distilled.
Four tests for any twin
-
Ask something new. Pick a situation the owner has never written about. Does it reason from their rules, fall back on generic advice, or say honestly that there is no recorded view?
-
Ask twice. Put the same question in different words on different days. The wording can change. The position should not.
-
Ask why. Does it point to specific principles and the owner's own words, or restate the answer?
-
Correct it, then come back in a week. Did the correction hold?
Run these on Imora as well. The Free tier is permanent, so there is no deadline on the test. The comparison page sets out where we think other tools are ahead.