How to Write Prompts for Realistic AI Humans
The specific prompt details that separate a convincing AI-generated person from one that reads as obviously synthetic.
Human figures are the single hardest thing for AI video models to render convincingly — faces, hands, and skin texture are where the "uncanny valley" shows up most obviously. Getting a realistic-looking person isn't about finding one magic prompt phrase; it's about layering several specific details that, together, push the output toward realism instead of the smooth, slightly-off default many models fall back to.
Skin Texture Is the Biggest Single Lever
Explicitly requesting realistic skin texture — pores, subtle imperfections, natural variation in tone — meaningfully changes output away from the overly smooth, plastic-looking default many models produce for human skin. Terms like "realistic skin texture," "natural skin detail," or referencing photography styles known for texture (like documentary or editorial photography) all push in this direction.
Describe Imperfection, Not Perfection
Ironically, asking for a "perfect," flawless face often produces something that reads as more artificial, because real human faces are asymmetrical and imperfect. Small details — a slight asymmetry, natural under-eye shadow, realistic hair flyaways — push a generation toward believability precisely because they mimic the kind of small imperfections a camera captures in real people but a model defaults away from unless specifically asked.
Lighting Reveals or Hides Artifacts
Harsh, flat, even lighting tends to expose whatever artifacts a model produces around the face — it has nowhere to hide. Slightly directional, soft lighting (like window light or golden hour) is more forgiving and tends to produce more convincing results, partly because it's a lighting condition the model has seen far more often in its training data specifically paired with realistic human portraits.
Camera Distance Matters More Than People Expect
Extreme close-ups on a face put the most demand on a model to get fine detail exactly right, and are where artifacts are most visible to a viewer. A medium shot or a slightly further framing gives the model more room for error that a viewer is less likely to consciously notice, even though the same underlying quality issues are technically present. If a specific close-up keeps producing artifacts no matter how the prompt is adjusted, pulling the camera back slightly is often a faster fix than continuing to tweak the wording.
Hands Are Still a Known Weak Point
Even with excellent face realism, hands remain one of the most commonly distorted elements across most current AI video tools. Where possible, frame shots to keep hands out of frame, partially obscured, or in motion (motion blur hides fine detail issues more forgivingly than a static, clearly visible hand). This is a known, common limitation across the industry, not a sign of doing something wrong in your prompt.
Clothing and Fabric Realism
The same realism principles that apply to skin also apply to clothing — describing fabric type and how it behaves ("heavy wool coat," "lightweight cotton shirt with natural wrinkles") produces more convincing results than leaving clothing generic and undescribed. Fabric that moves and folds naturally is one of the visual cues that reads as "real" to a viewer even below the level of conscious awareness.
Avoid overly pristine, unwrinkled clothing descriptions unless the scene specifically calls for it (a formal event, a brand-new outfit) — real clothing in real footage almost always has some natural wrinkling and movement, and requesting that explicitly tends to push a generation away from the artificially smooth default.
Ethnicity, Age, and Representation Details
Being specific about a character's ethnicity, approximate age, and other demographic details isn't just about representation — it also meaningfully narrows the model's guessing space and tends to produce more consistent, higher-quality results than a vague, unspecified description that forces the model to default to whatever's most statistically common in its training data.
Write these details respectfully and specifically, the same way a casting description in a real film production would, rather than vaguely — this level of specificity is one of the most underused levers beginners have available for improving both realism and consistency in human-focused AI video.
Eye Detail and Gaze Direction
Eyes are one of the most scrutinized details in any human-focused shot, and explicitly describing eye color, expression, and gaze direction (looking directly at camera, looking off to the side, eyes closed) gives the model clearer guidance than leaving this detail unspecified. A mismatched or unfocused gaze is a subtle but common giveaway in AI-generated human faces, and directly addressing it in the prompt measurably reduces how often this issue appears.
Referencing Real Photography Styles for Believability
Referencing a specific, well-known photography style — documentary photography, editorial portrait photography, candid street photography — carries strong implicit realism cues that the model has learned from extensive training data tagged with those exact terms. This is often more effective than trying to describe realism directly ("make this look real"), since style references pack in dozens of specific visual conventions (natural lighting, unposed expression, genuine texture) that would take many more words to describe individually.
Balancing Realism With Your Channel's Actual Style
Not every channel benefits from maximum photorealism — a stylized or slightly cinematic look can be a deliberate creative choice rather than a limitation. Decide early whether your channel's identity leans toward strict realism or a more stylized look, and apply these realism techniques selectively based on that choice rather than defaulting to maximum realism for every single shot regardless of context.
Reviewing Realistic Generations to Reverse-Engineer What Worked
Whenever a generation comes out looking particularly convincing, take the time to analyze exactly which prompt elements likely contributed — the specific skin texture phrasing, the lighting choice, the camera distance — rather than moving on immediately to the next shot. This reverse-engineering habit, done consistently, builds a much deeper practical understanding of realism than reading general guidance alone.
Patience With Iteration on Human-Focused Shots
Human-focused shots typically require more generation attempts to get right than environmental or object-focused shots, simply because there's more fine detail for a model to potentially get wrong. Budgeting extra generation attempts specifically for shots featuring people — rather than expecting the same one-or-two-attempt success rate you might get on a simpler landscape shot — sets more realistic expectations for the actual production time a human-focused project requires.
Want the exact prompts behind our AI videos too?
Browse All Tutorials