Journal · Methodology · plain words
Why average is fifty, not seven out of ten
BecomeTen puts an average, healthy, groomed adult at 50 to 60 because a scale where everyone is a 7 cannot tell anyone anything. The eight tiers, five weighted areas and per-area potential caps anchor the number, so a small change is visible and a high score is earned. Average is a description, not an insult.
Why does everyone get a 7 out of 10?
Ask a friend to rate you out of ten and you will get a 7. Ask most face-rating apps and you will get something between 7 and 8.5. This is not because most people are a 7. It is because a rater who hands out a 5 has to defend it, and a rater who hands out a 7 does not. Politeness compresses the scale from the bottom, and the top compresses itself because nobody wants to give a 10 to someone they know. What is left is a band between 6.5 and 8 where nearly everyone sits and nothing is distinguishable.
An app has an extra reason to inflate: a flattering number is shared and a plain one is not. So the drift is structural, and it means a 7 from most sources carries almost no information. You cannot tell a 7 from a 7.5, you cannot tell whether a change moved you, and you cannot tell whether the same face would get a 7 tomorrow. The article on why PSL ratings vary covers the forum version of the same problem.
What anchors the BecomeTen scale?
A number is only strict if something outside the number holds it in place. BecomeTen uses three anchors, and all three are written down in the methodology rather than left to the model's mood.
The first is the eight tiers. Each covers a fixed band of the 0 to 100 scale and has a written description of what a face in that band looks like. An average, healthy, groomed adult is the definition of the Normie tier, 50 to 64, and the rubric says so in words before any photo is scored. The tiers above it are defined by what is rare: clearly above the midpoint, striking structure, near the top in most areas, the ceiling. The tiers page has each definition.
| Anchor | What it fixes | What it prevents |
|---|---|---|
| Eight tiers with written definitions | Where average sits, and what the bands above it require | Drift toward a polite 7 for everyone |
| Five weighted areas, 30+ traits | The score is built up from parts, not guessed as a whole | A halo from one strong feature or one flattering photo |
| Per-area potential caps | How far each area can move without surgery | A potential score that assumes bone will change |
The second anchor is the five weighted areas: harmony, eye area, jaw and chin, skin, hair, each built from its own traits and each weighted by how much it drives how a face is read, with harmony carrying the most. The overall score is assembled from those parts, which means the model cannot simply decide a face is a 72; it has to have found traits that add up to one. The third is the set of potential caps: harmony up to 6, eye area up to 10, jaw and chin up to 14, skin up to 25, hair up to 20. They bound the second number, potential, so that it can never promise more than routines change.
Why does a strict scale make a 6 mean something?
On a compressed scale a six-point gap is the whole visible range: it separates everyone from everyone. On a scale where average is 50 to 60, six points is a real and specific distance, roughly the difference between a face with one clearly weak area and the same face with that area brought to average. It is large enough to be a genuine change and small enough to be reachable, which is the size a useful number has to be.
The same stretch is what makes the top of the scale honest. Because the middle is genuinely the middle, a score above 80 has room to mean what the tier says: striking structure with few weak points. It is reserved, and it is rarely assigned, and that is the point of having it. A scale where 8 is common has no way to describe someone who is actually unusual.
It also makes a low number survivable. A 42 on BecomeTen says that one or two areas are pulling the rest down, and the report says which; in most photos those areas are workable ones. On a ten-point scale the same face is a 6.5 with no explanation, which sounds kinder and is less useful.
How does the PSL equivalent fit?
The PSL scale is the forum's 1 to 8 ladder, and it is strict in the same direction: the forums put an average face around 4 to 5 and reserve 7 and above for a handful of people. BecomeTen shows a PSL equivalent next to the score because many readers already think in that ladder. It is a translation of the 0 to 100 score, computed from it, not a separate rating, and the PSL score page sets out the mapping.
Two consequences. First, an average adult at 50 to 60 lands at about PSL 4 to 5, which is what the forums would also call average, so the two scales agree where it matters most. Second, because it is a translation, it inherits the anchors: nobody gets a PSL 7 from a flattering photo, because nobody gets the 0 to 100 score that maps to it without the traits that justify it.
Is average an insult?
No, and the research makes the point better than we can. Average is what most faces are, by definition; a scale on which most people do not land in the middle is a scale that has stopped describing anything. On BecomeTen, Normie is the largest tier because it should be.
So “average” in the score means the whole is in the middle of the scale, and it usually means the proportions are close to the population mean, which raters like. It does not mean unremarkable to look at, and it does not mean nothing is good. The tier describes a healthy, groomed adult whose face has no serious weak point; that is a description most people would accept of someone else, and struggle to accept of themselves.
How does a strict scale make progress visible?
Because the score is built from traits and stretched across the full range, the things that actually change a face show up as changes in the number. Clearer skin over three months of a routine, a leaner jaw after body-fat loss, a denser hairline after a course of treatment: each moves its area, and the area moves the score by an amount the caps bound. On a compressed scale those changes disappear into the half-point of noise between 7 and 7.5. The how AI face rating works article walks through the assembly step by step.
The strictness is also what makes a re-scan worth doing. If the same photo conditions give the same score, a change in the score is a change in the face, and the tracking guide explains how to hold the conditions steady. A generous rater cannot offer that, because its number was never attached to anything in the first place.
Rate my face
Score, tier and PSL equivalent, calibrated so that average means average.
Questions
Is a 55 on BecomeTen the same as a 5.5 out of 10?
No. A 55 sits in the Normie tier, which is defined as an average, healthy, groomed adult, and that is roughly a 7 on the polite ten-point scale most people use. The scales are stretched differently: BecomeTen spreads faces across the full range so the middle is genuinely the middle, while a ten-point rating squeezes nearly everyone into 6.5 to 8.
Why not just calibrate the scale so average is 70?
Because then the top thirty points would have to hold everyone above average, and a scale with no room at the top cannot distinguish clearly above average from striking. Putting average at 50 to 60 leaves space on both sides, so a weak area and a strong one each move the number by an amount you can see. The tier names, not the digits, carry the meaning.
Does a strict scale mean nobody scores above 80?
It means few do. The High tier and above are defined by striking structure with few weak points, and that is rare in any population. People do land there, usually with strong harmony and eye area, and the report shows the traits that put them there. A strict scale is what makes such a score credible rather than a compliment.
Why did my friend rate me higher than BecomeTen?
Because your friend was being polite, and because a friend rates a person while the model rates a photo against a rubric. Both are honest in their own way. Humans compress the scale from the bottom out of kindness and from the top out of reluctance; a rubric with written tiers does neither, so its numbers are lower and mean more.