Journal · Scales · context
Why two raters give the same face a different PSL
PSL ratings vary because the scale has no rubric: each rater brings their own reference faces, their own weighting of eyes, jaw and skin, and their own mood, and the photo adds angle, light and lens distance on top. A spread of one to one and a half points between raters on one person is normal.
Why is there nothing to anchor a PSL rater?
The PSL scale is a 1–8 ladder with a firm midpoint and a strict ceiling, and nothing else. There is no list of traits a rater must look at, no weighting between them, no reference set of faces pinned to each rung and no procedure for the photo. Two people who both know the ladder well can therefore apply it honestly and land a point and a half apart, and neither is doing it wrong, because there is no way to do it right.
Start with reference faces. Every rater carries a private gallery of faces they have already placed, and a new face is rated by comparison with that gallery. Someone whose gallery is built from model portfolios has a 6 that sits higher than someone whose gallery is built from a university campus. Neither gallery is wrong; they are simply different rulers, and a rating is only as transferable as the ruler behind it.
Then weighting. One rater treats the eye area as most of the score, another the jaw and chin, a third cannot see past skin. Forum culture leans structural, so a face with a strong gonial angle and poor skin tends to rate higher there than the same face would with a general audience, but even inside a single thread the weightings are all over the place. A face that is uneven across areas, strong in one and weak in another, gets the widest spread, because it lands wherever each rater's weighting says it should.
And then mood and motive. Raters rate more harshly after a run of striking faces and more kindly after a run of ordinary ones, an ordinary contrast effect. Some rate to be accurate, some to be kind, some to be cruel, and some to signal that they have high standards. A rating given to a stranger in an anonymous thread carries every one of those motives, and the number alone does not tell you which.
How much does the photo move a PSL rating?
Almost every PSL rating is a rating of a photograph, and a photograph is a set of choices. Angle first: a chin tilted a little down sharpens the jaw and narrows the lower face, tilted up it does the reverse and shows the nostrils; a three-quarter turn hides asymmetry that a straight-on view shows. Light second: soft front light flattens skin texture and fills the under-eye, hard overhead light does the opposite and can take a healthy face into tired-looking in one frame.
Lens distance is the one people underestimate. A phone held at arm's length puts the lens close enough that the nose and the middle of the face enlarge relative to the ears and the jaw, which reads as a weaker jaw and a longer midface than the person has. Step back and zoom, or have someone else take it from a couple of metres, and the same face reads more balanced. Expression finishes the job: a smile raises the cheeks, narrows the eyes and changes the read of the whole eye area, and a slightly open mouth lengthens the lower third.
Put the four together and the same face can land a full point apart on the ladder between a good photo and a bad one, before any rater has disagreed with any other. That is why the first thing worth doing about a rating you do not like is to look at the photo, and why photo quality, light, angle and lens is the most useful page on this site for anyone who has been rated online. How to take a photo for a face rating turns it into a checklist.
Are half-points the ladder's resolution limit?
Yes. The community works in half-points and often in tenths, and argues fiercely about them, but the arithmetic does not support the precision. If rater-to-rater spread on one photo is around a point, and photo-to-photo spread on one face is around a point, then a half-point difference between two ratings is well inside the noise. It is the smallest step the ladder offers, not the smallest difference it can detect.
This matters most in the range where most people sit. A 4.5 and a 5 are a different tier in the community's vocabulary and the same face in practice. When someone reports moving from 4.5 to 5 after a routine, what usually happened is a better photo, a kinder thread, or both. A real change shows up as a shift that survives several photos and several raters, which is a much higher bar than a half-point in one thread.
Does PSL's strictness make it more accurate than a 1–10?
Stricter, yes; more accurate, no. The everyday 1–10 fails because almost everyone is rated a 6 to 8, so the scale has no room to tell people apart. PSL pulls the ceiling down to 8 and puts the ordinary face at 4, which spreads real faces across more of the scale and makes a 6 mean something. That is a genuine improvement in range. It does nothing about variance, because the strictness lives in the rater's head and every rater's head is different.
The strictness does explain one thing people find hard to accept: an average, healthy, groomed adult is a 4 to 5 on PSL, and on BecomeTen lands around 50 to 60, which the PSL score page converts to the same range. That is not a harsh reading of the person. It is what the midpoint means. A scale that told everyone they were above average would be telling them nothing, and the people who search for PSL are usually the first to notice when a rating is flattering them.
What do observers actually agree on?
The variance is real, but so is the agreement underneath it, and perception research is clear about where each lives. Observers agree, far above chance, on which of two faces is more attractive, and the cues that carry that agreement are broad ones: symmetry, closeness to the population average in shape, visible skin health and cues of body fat. They disagree on the fine ranking inside a tier, on how much a single trait should count, and on where the boundaries between tiers fall, which is exactly the part of the job the PSL ladder asks them to do with half-point precision.
So the honest summary is that raters agree on direction and disagree on the number. A face that most raters place above the midpoint really is above it; whether it is a 5 or a 5.5 is not a fact about the face at all.
How BecomeTen handles the variance
A crowd's spread cannot be fixed by adding more crowd. What can be fixed is the rater side and the photo side, and BecomeTen does both. One rubric of 30 or more traits is applied to every face in the same order with the same weightings, so the reference faces, the weighting and the mood are constant; the rating is repeatable, cannot be argued upward and does not depend on who is in the thread. The score is produced on 0–100 and shown with a PSL equivalent beside it; the conversion is a fixed table on the PSL score page, not a second opinion.
The photo side is handled by grading the photo before rating the face. Angle, light, lens distance and expression are each checked, and a photo that fails the check is flagged so the number is read as an estimate with a wide band rather than a verdict. Because capture still moves a clean photo by a point or two on the 0–100, a difference within two points between scans is shown as noise, neither progress nor decline. On the PSL side that is a fraction of a half-point, which is the same conclusion the ladder's own arithmetic reaches.
The practical advice follows from all of this. Rate the same clean photo setup every time. Read a one-point PSL difference between two raters, or between a forum and an app, as ordinary spread. Watch for a change that survives several photos over weeks, which how to track progress shows how to do, and spend the effort on the workable areas, skin, hair and body composition, where looksmaxxing without the bro-science actually moves a rating.
Rate my face
One rubric, one graded photo, the PSL equivalent shown beside the score. Free, no sign-up.
Questions
Why does BecomeTen give a lower PSL than a forum did?
Usually because the forum rater was generous, the photo was flattering, or both. BecomeTen is calibrated so that an average groomed adult lands at a 4 to 5, applies the same rubric to every face and grades the photo for angle, light and lens before rating it. A half-point to a point between an app and a stranger is inside the normal spread between any two raters; a wider gap almost always traces back to the photo.
Are PSL half-points meaningful?
Only as a community convention. A half-point of PSL is about six points on BecomeTen's 0–100, and between two photos of the same face that much can come from angle and light alone, before two raters have disagreed at all. Treat a half-point as the resolution limit of the ladder, the smallest step it offers, not as a real difference between two people or between two versions of yourself.
How much can a PSL rating change between two photos?
About a full point, in either direction, from capture alone. A chin tilted down, a lens too close, hard overhead light and a wide smile each shift how a face reads, and together they can take the same person from a 4 to a 5 or from a 5 to a 4 without anything about the face changing. Two clean photos taken the same way that agree with each other are worth more than any single rating.
Do PSL raters agree with each other at all?
On direction, yes; on the number, less than they think. Perception research finds observers agree well above chance on which of two faces is more attractive, driven by symmetry, averageness, skin and body-fat cues. Raters disagree on how much each trait counts and where the tier boundaries fall, so a spread of a point to a point and a half on one face in one thread is ordinary and does not mean anyone is lying.
Sources
- 1.Human (Homo sapiens) facial attractiveness and sexual selection: the role of symmetry and averageness — Journal of Comparative Psychology (1994)(Opens in a new window)
- 2.Facial attractiveness: evolutionary based research — Philosophical Transactions of the Royal Society B (2011)(Opens in a new window)