Most people generating AI images have quietly given up on one thing: getting the same person to appear twice. You produce a great character, then the next generation gives you a stranger with a similar haircut. So you settle. You use one image, or you crop tight, or you switch to stock.
That limitation ended sometime in the last year, and the fix is not a better prompt. It is a reference workflow. Below is how character consistency actually works in 2026, which two methods are worth your time, and the exact steps to lock a face down across a whole campaign.
What character consistency actually means in AI image generation
Character consistency is the ability to place the same recognisable subject, the same face, body and wardrobe, into new scenes, poses and lighting without the identity drifting. It is a reference problem, not a prompting problem. Text alone cannot describe a face precisely enough, so consistency comes from feeding the model an image of the subject on every generation.
This matters because it changes what AI images are for. A single hero image is decoration. A consistent character is an asset: you can build a six-slide carousel, a product explainer, a training deck or a month of social posts around one recognisable person. That is the difference between using AI for pictures and using AI for campaigns.
Why your character drifts between generations
Drift happens because image models sample fresh noise every run. Your prompt narrows the space of possible faces, but a text description like "woman in her thirties, shoulder-length dark hair" still permits millions of distinct faces. Without a visual anchor, each generation independently picks one.
Three things make drift worse in practice.
Prompt-only identity descriptions. Adding more adjectives feels like it should help, and it does narrow things slightly, but the returns collapse fast. Twenty words about a face buys you far less identity lock than one reference image.
Switching models mid-project. Every image model has a house look. The same reference and the same prompt will render differently across Midjourney, Nano Banana and a Flux checkpoint. Pick one model per character and stay there.
Letting the scene overwhelm the subject. Long, dense scene descriptions compete with identity for the model's attention. If your prompt spends forty words on a rain-soaked Central street at dusk and six words on your character, the street wins.
Method 1: Nano Banana Pro multi-reference, the fastest reliable route
Google's Nano Banana Pro accepts multiple reference images in one request and reconstructs the subject's identity from them. The practical recipe is to upload three to five images of the same person from different angles, then describe only the new scene. It handles up to five distinct reference subjects in a single frame and is accessible free through Google AI Studio.
This is the method to start with if you have never done reference-based generation, for three reasons. It needs no parameters to learn. It tolerates imperfect references, phone photos work. And because it accepts several subjects at once, you can put two consistent characters in the same shot, which is where most single-reference tools fall apart.
The angle spread is the part people skip. Three images shot from the same front-facing angle give the model one view of a face and nothing about its structure. Front, three-quarter and profile give it geometry, and geometry is what survives a change of pose.
Method 2: Midjourney Omni Reference and the weight dial
Midjourney handles identity through Omni Reference, invoked with the --oref parameter pointing at a reference image URL, plus --ow to set how hard the model clings to that reference. The weight is the whole craft: low values let the scene and style dominate, high values preserve the face but flatten your stylistic control.
In practice a moderate weight is the working default. Push it up when the face is the point, a portrait, a talking-head thumbnail, a headshot in a new outfit. Pull it down when you want your character rendered in a strong illustrative style, because a high weight will drag the output back toward the photographic look of your reference and fight the style you asked for.
Midjourney remains the better choice when the aesthetic matters more than the pixel-accuracy of the face. Nano Banana is the better choice for controllable editing, multi-image workflows and production work where the same person must be unmistakable across twenty assets. That is the honest split, and it means the answer to "which one" is usually "both, for different jobs."
The canon-image workflow that ties both methods together
Whichever tool you use, the discipline is the same: designate one canonical image and feed it back on every single generation. Never generate from a generation. If you use output number four as the reference for number five, small errors compound and by image twelve you have a different person.
Here is the workflow that holds up across a real campaign.
Step 1. Generate or shoot a character sheet: one subject, neutral lighting, plain background, three angles, neutral expression. Neutral is deliberate, a strong expression bakes itself into every later image.
Step 2. Pick the single best front-facing frame and name it canon. Save the other angles as supporting references.
Step 3. Write your scene prompt with identity described briefly and scene described fully. The reference carries the face; the words carry the world.
Step 4. Generate in small batches, four at a time, and reject anything with drift immediately rather than trying to fix it in a follow-up edit.
Step 5. Keep a plain text file of the exact prompt, model, parameters and reference filenames. Six weeks later, when the client wants one more asset, this file is the only thing that will let you match the set.
Where character consistency still breaks down
Reference-based generation is reliable now, but not universally. Four failure modes are worth knowing before you promise a client a series. Hands and fine jewellery still drift, extreme camera angles lose identity, text on clothing garbles, and children's faces are noticeably less stable than adults'.
Wardrobe is not identity. Reference images lock the face far more strongly than the outfit. If a specific jacket or uniform matters, describe it explicitly in text on every generation and expect to reject a portion of outputs.
Extreme angles. A shot from directly overhead or from below the chin gives the model almost no facial geometry to match against. Consistency degrades sharply. Design your shot list around eye-level to three-quarter angles.
Groups. Two consistent characters in one frame is achievable with multi-reference. Four is not, reliably. Compose groups as separate generations where possible, or accept that background figures will not hold identity.
Style transfer plus identity. Asking for a heavily stylised render, watercolour, anime, 1970s film grain, while also demanding a precise face is the hardest ask in this space. Something gives. Decide in advance which one matters more, and set your reference weight accordingly.
How to test whether your reference is actually working
Before you build a campaign on a character, spend five minutes validating the reference. The test is simple: generate the same subject in four deliberately different conditions and check whether a stranger would identify them as one person. If three of four hold, your reference is production-ready.
Use these four checks in order. First, a pose change with identical lighting, this isolates whether the model has facial geometry or only a memorised photo. Second, a lighting change with an identical pose, warm indoor to overcast outdoor, which exposes references that leaned on a specific colour cast. Third, a distance change, from close portrait to mid-shot, since some references only survive at the framing they were shot in. Fourth, a wardrobe change, which is the one most likely to fail and the most useful to know about early.
Score it honestly and write the result down next to the reference filename. A reference that passes three of four is worth keeping; one that passes two will cost you more time in rejected generations than it saves. Regenerating a character sheet is cheap. Discovering the problem on asset seventeen is not.
Try this now: the character sheet prompt
Copy this into your image tool of choice to generate a reusable character sheet. Replace the bracketed description once, then never change it again for this character.
Prompt:
Character reference sheet, three views of the same person side by side on a plain light grey seamless background: front view, three-quarter view, and full profile view. Subject: [a Hong Kong Chinese woman in her early thirties, shoulder-length straight black hair parted slightly off-centre, minimal makeup, wearing a plain charcoal crew-neck top]. Neutral relaxed expression, mouth closed, eyes open and looking at camera in the front view. Even soft studio lighting, no harsh shadows, no colour cast. Sharp focus on facial features. Consistent identical facial structure across all three views. Photographic, natural skin texture, no retouching. 16:9 wide composition.
Then, for every scene image after that, use this template and attach the canon frame as your reference:
Scene prompt template:
[Reference: canon.jpg] The same woman from the reference image, now [seated at a cafe table in Sheung Wan reviewing a printed report, mid-morning daylight through a window to her left, shallow depth of field, warm neutral colour grade, candid documentary photography style]. Keep facial identity exactly matching the reference. Natural pose, no eye contact with camera.
Run that pair once and you will see the difference within four generations. If you want the reasoning behind why the brief identity description outperforms a long one, our piece on what replaced "act as an expert" prompting covers the same principle in text work.
The takeaway
Character consistency stopped being a prompting skill and became a filing skill. The people producing coherent AI image sets in 2026 are not writing cleverer descriptions. They are keeping one canon image, one model, one parameter set, and a text file that records all three. That is unglamorous, and it is the whole trick.
Start with a character sheet this afternoon. Three angles, neutral light, one subject. Everything else in this article depends on that one file existing, and it takes twenty minutes to make. Tooling changes fast in this field, but the reference discipline has outlasted three generations of image models and will outlast the next.
We understand AI. We understand you better. With UD by your side, AI doesn't feel cold.
Reviewed by the UD AI team.
Turn One Technique Into a Working Pipeline
A reference workflow is the easy part. Making it run reliably across a whole marketing team, with shared assets, approved characters and a repeatable brief, is where most teams stall. We'll walk you through every step, from tool selection and reference libraries to workflow design and deployment.