How big a face has to be in a thumbnail
Every guide tells you to use a face. None of them says how big, which is the only part that is decidable — a thumbnail is rendered anywhere from 0.28× down to 0.13×, and a face that reads at one scale is a smudge at the other.
“Use a face” is not a size
It is the most repeated piece of thumbnail advice and the least actionable, because it stops exactly where the decision starts. A face at what size? You are designing on 1280 × 720 and the interface will not show you that. It shows a 346-pixel-wide card on a desktop home feed, a 360-pixel row in search, and a 168-pixel item in the suggested column beside a playing video.
That is a reduction to 0.27× on a card and 0.13× in the suggested column — the same file, rendered at a seventh of its linear size. Face size is the one property of a thumbnail that gets destroyed by that reduction and can still be fixed for free while you are designing, because it costs nothing to make a head bigger. YouTube's own thumbnail and title tips say to show a face too, and attach no number to it — which is the gap this page exists to fill.
What has to survive is not the head
A head is a silhouette. At almost any size it tells you a person is in the frame, which is worth something but is not the reason anyone recommends faces. The reason is expression — surprise, doubt, effort, disgust — and expression lives in a band across the eyes and the mouth.
That band is about a third of the head. In the three real thumbnails I measured on a percentage grid it was 34%, 32% and 40% of head height, so a third is the working figure. Which means that whatever you think you are putting in the frame, the part that carries the message is a third as tall as you think it is.
A yardstick that is already on the page
Rather than claim a threshold in pixels for when a human face reads — I have not tested that and neither has anyone who tells you 20% or 30% — use a comparison you can check yourself, on the same screen, in the same glance.
Next to every thumbnail YouTube renders your title, and one line of it is 20 pixels tall on a card, 23 in a search result, 18 in the desktop suggested column and 16 on mobile suggested. If the eyes-and-mouth band of your face is shorter than one line of the title beside it, it is not competing for attention with that title. It is texture.
So here is how much of the frame the head has to fill
One line of title as the target for the band, and a band that is a third of the head, gives a required head height per surface. These are the rendered sizes of the mock in this tool, which follows YouTube's own layout:
| Surface | Rendered thumbnail | One line of title | Head must fill |
|---|---|---|---|
| Search result | 360 × 203 px | 23 px | 34% of frame height |
| Desktop home card | 346 × 195 px | 20 px | 31% |
| Mobile feed card | 370 × 208 px | 20 px | 29% |
| Mobile suggested list | 160 × 90 px | 16 px | 53% |
| Desktop suggested column | 168 × 94 px | 18 px | 57% |
The narrowest surface governs, and it is the same one that governs title length: the suggested column beside a playing video, where the viewer is choosing what to watch next. On a card a third of the frame is enough. There, it takes over half.
Three real thumbnails, read off a grid
To keep this out of the realm of assertion I took three thumbnails from the packs in this tool — a business talking head, a political split screen, a fitness demo — overlaid a grid at 5% intervals and read the head box and the eye-to-mouth band off it. Then multiplied by the rendered heights above.
| Thumbnail | Head | Band | Where the band drops under one line |
|---|---|---|---|
| Talking head, face on the right third | 65% of frame height | 22% | nowhere — 21 px even in the suggested column |
| Split screen, two faces | 53% and 40% | 17% and 13% | both suggested columns (16 px against an 18 px line) |
| Demo shot, face above the subject | 25% | 10% | everywhere but a feed card, where it lands exactly on one line (20 px) |
What that means in practice
The talking head is the only one of the three that works on every surface, and it looks absurd in the editor: a head occupying two thirds of the frame height, cropped at the top, taking a third of the width for itself. That is not a stylistic tic of big channels. It is what the arithmetic asks for.
The split screen is the interesting case, because it fails in exactly one place — the suggested column — and that is not a corner case. It is where a viewer is deciding what to watch after the video they are already watching, and it is the surface where nothing else about your thumbnail is legible either.
The demo shot is the common mistake and the most forgivable one, because the face is not meant to be the subject: the whiteboard is. The problem is that the face is still spending a quarter of the frame while carrying nothing at card size — ten percent of the height is nine pixels of expression in a suggested column. If the subject is the subject, that quarter of the frame is better spent on the subject.
None of this says a face must fill a third of the frame. It says a face that is *meant to carry expression* must, and a face below that size is decoration you are paying frame area for. Deciding it is decoration is a legitimate choice; not deciding is the expensive one.
Two faces is not two chances
A split frame is the standard way to signal a confrontation, a comparison, a before and after, and it is worth knowing what it costs. Each face gets half the width, so a head that would have been 65% of the frame height at full width is typically 40–55% in a half — which is what I measured in the political one: 53% and 40%.
The larger of the two survives on cards and dies in the suggested columns; the smaller dies earlier. So a split frame effectively trades the narrowest surface away in exchange for a relationship between two subjects. Sometimes that relationship *is* the video and the trade is right. What it is not is a way to double the amount of face in the frame.
Where the face can actually go
Size decided, position has two constraints, and neither is about composition.
The interface draws over the frame: the duration badge in the bottom-right corner, badges in the bottom-left, the watched-progress bar along the bottom edge. A face whose mouth or chin sits in the bottom-right is a face with a black pill over it on a good part of its impressions. The zones are laid out in the safe zone guide, with a template you can drop over your design.
And a Short is a different image entirely, not a crop of this one — YouTube serves a genuinely vertical still there, which is why the Shorts guide treats it separately. If the same face is meant to work in both formats, it has to be composed twice, and the vertical version has the easier job: a 9:16 frame gives a head far more room to be large.
A big face leaves less room for text, and that is fine
The objection to a head at a third of the frame is always the same: there is no room left for the words. Correct, and it is not a loss, because the words that matter are not in the image.
The title beside the thumbnail is rendered by YouTube at a size it controls and is guaranteed legible — that is where numbers, names, models and years belong. The image carries one wordless idea, and a face at working size is one of the better ways to carry it. That division is the subject of the guide on the thumbnail–title pair, and it is what makes room for a large face rather than competing with it.
If text still has to be in the image, it is competing with the face for the same frame area and both shrink. Three or four words at a size that survives the reduction is the ceiling; the guide on text size has the measurement for that half.
When a face is the wrong call
Two cases where the arithmetic argues against a face rather than for a bigger one.
When the subject is an object — a product, a chart, a place, a screen — a face at working size is a third of the frame spent on a person the viewer does not yet recognise. Recognition is what makes a face valuable on a channel with an audience, and a channel without one gets only the expression, which has to be strong enough to be worth that area.
And when the expression is neutral. A calm, well-lit, professional face at 65% of the frame is a large photograph of a stranger. The size rule is a floor for legibility, not a claim that any face at that size works; it means that if the expression is worth showing, this is what it costs to show it.
The one-minute check
Everything above is an argument about a reduction, and a reduction is something you can just look at. Put the real file into a real feed and go to the suggested column — the narrowest surface, the one that governs — and see whether the face is still doing a job there or has become a shape.
That is what the preview tool is for: the same image at all five rendered sizes, side by side with the competitors it will actually appear beside. A face that survives the suggested column survives everywhere else by definition, so it is the only view you strictly need to check.
The questions everyone asks
How big should a face be in a YouTube thumbnail?
Big enough that the eyes-to-mouth band is at least as tall as one line of the title beside it. That works out to about a third of the frame height on a feed card and over half in the suggested column, which is the surface that governs.
Do thumbnails with faces perform better?
A face is a strong way to carry one wordless idea, but only at a size where the expression survives the reduction. Below that size you have paid frame area for a silhouette, which is why “use a face” without a size is not useful advice.
Is a face necessary on every thumbnail?
No. When the subject is an object, a face at working size costs a third of the frame for a person the viewer may not recognise. On a channel with an audience recognition pays that back; on a new channel only the expression does.
Can I put two faces in one thumbnail?
Yes, and it halves the width available to each, so heads land around 40–55% of frame height instead of 65%. Measured on a real split screen, the larger face survives on cards and both faces fall below one line of title in the suggested columns.
Where should the face go in the frame?
Anywhere except the bottom-right corner, where the duration badge is drawn, and off the bottom edge, where the watched-progress bar sits. A face's mouth and chin in the bottom-right is the common version of this mistake.