Journal · Methodology · plain words
How AI face rating works, step by step
An AI face rater sends your photo to a vision model with a fixed rubric and turns its structured opinion into a score. In BecomeTen the photo is graded for quality first, the headline is recomputed on the server from weighted areas, potential is capped in code, and every citation in the plan is checked against a curated library.

What happens to the photo, step by step
The pipeline is short, and every stage exists to stop the model from being trusted more than it deserves.
- You upload one photo. For a free scan it is processed for the rating and then deleted.
- The photo is graded for capture quality: lighting, angle and sharpness.
- A vision model is asked for a structured opinion against a fixed rubric: five areas, 30+ traits, each trait as a directional band.
- The server recomputes the headline score from the five weighted area scores rather than trusting the model's own arithmetic.
- The server caps each area's potential at a fixed non-surgical ceiling written in code.
- A second prompt writes the plan. It may only cite from a curated library of publications, and every id it returns is checked.
- You see the score and tier free; the trait breakdown, plan and potential photo are behind the one-time report.
The capture-quality grade
Before anything is scored, the photo gets three grades: lighting is good, uneven or poor; angle is frontal, slightly off or off; sharpness is good, soft or blurry. A photo that fails a grade is still rated, because refusing would be worse than a noisy number, but it is flagged so you know the rating is less precise. On a re-scan, a failed grade on either photo means the comparison with your baseline is withheld. What each of those failures does to the score is covered in photo quality: light, angle and lens.
A structured opinion against a fixed rubric
The word to hold onto is opinion. The model is not measuring anything; it is looking at a two-dimensional image and answering a fixed set of questions about it. The rubric has five areas — harmony, eye area, jaw and chin, skin and hair — and each area has a list of traits: facial thirds, fifths, symmetry and midface ratio under harmony; canthal tilt, eye spacing, under-eye and brow under eye area; gonial angle, jawline definition, submental definition and chin projection under jaw and chin, and so on.
Every trait comes back as one of four bands: elite, strong, average or weak. Not a number. The rubric forbids the model from stating a degree or a millimetre for a structural trait, because that cannot be read from a selfie and pretending otherwise is false precision. The calibration instruction is strict on purpose: an average, healthy, groomed person should land around 50–60, and above 80 is reserved for genuinely striking structure.
Rate my face
One photo, about 60 seconds — score and tier free, no card.
Why the server recomputes the headline
A language model is good at judgement and unreliable at arithmetic. If the model were allowed to return its own overall score, it could hand you a 71 alongside a breakdown that averages to 63, and the first person to check would be right to distrust the whole report. So the model's headline is discarded. The server takes the five area scores, applies fixed weights — harmony carries the most, then eye area, then jaw and chin, then skin, then hair — and rebuilds the overall from those. The same happens for potential.
The weights are a product decision fixed in code; the caps and the strict calibration they feed are public in the methodology. Structure carries most of the weight because a looksmaxxing rating is mostly about structure; skin and hair are weighted lower because they are the most changeable, not because they matter least.
Potential is capped in code, not requested in the prompt
Potential is what is reachable without surgery. Left alone, a model asked for it will promise thirty points on bone. So the ceiling is enforced after the model answers: harmony can rise by at most 6 points, eye area 10, jaw and chin 14, skin 25, hair 20. Anything above the ceiling is clamped, and a potential below the current score is raised to match it. Harmony gets the least because bone geometry does not change; skin gets the most because it is where a routine with real evidence does the most visible work.
Each weak point is also tagged workable or structural. For a structural trait the plan may not promise a fix; it may only address how the trait reads — through composition, posture, hair framing and grooming. That tag is what separates a rating from a sales funnel.
The plan may only cite the library, and the server checks
A model asked to cite a paper will sometimes produce a plausible reference that does not exist. BecomeTen removes the option. The plan prompt is given a curated library of real publications — each hand-picked for an intervention the plan actually recommends, and each verified against its PubMed record — and may attach sources only by id from that list. When the plan comes back, the server resolves every id against the library and drops any it does not recognise. A fabricated reference cannot reach you.
Are AI face raters accurate? The honest limits
Accurate against what? There is no ground truth for attractiveness, so no rater can be validated the way a thermometer can. What can be said is narrower: the score is a consistent application of a fixed rubric, calibrated to a stated average, and it is an estimate from one photo, not a measurement.
- It sees one two-dimensional photo. Depth, profile and projection are inferred, not seen, which is why structural traits are bands and never numbers.
- The calibration is a choice. Someone else could set the average at 65 and call the same face something different; the number only means something on BecomeTen's own scale.
- The rubric is not adjusted for ethnicity or gender identity. The score describes the photo against one fixed rubric and that is a deliberate, stated limitation.
- It does not compare you to other users, and the percentile shown is a position on the scale's own curve, not a ranking of people.
Consistency between scans
Two scans of the same face in the same session will not always return the same number. The model's estimate moves a little with light, angle and focus, and a little on its own. BecomeTen counts a change within two points as noise and shows it neutrally. On a re-scan the model sees both photos together and may report only visible differences; unchanged and regressed are ordinary answers, and it may never report a change in bone between scans.
That is the practical meaning of consistency here: not that every scan matches to the point, but that the same photo conditions give the same band on each trait and the same tier on the ladder, and that a real change over weeks stands out from the noise. The rest of what the rater refuses to do — surgery, tissue damage, demeaning labels — is on the AI face rating page.
Questions
Which AI model rates the photo?
A current multimodal vision-language model, called with a fixed system prompt that defines the rubric, the bands and the calibration. The model is a component, not the product: the grading, weights, caps and citation checks all happen outside it, on the server, and are the parts that make the number trustworthy.
Why does the model return bands instead of numbers for traits?
Because a number would be false precision. A gonial angle or a canthal tilt cannot be read to a degree from a single front-facing selfie, and any tool that claims to is guessing with decimals. A band — elite, strong, average, weak — says what the photo can support and nothing more.
Could the model be biased toward certain faces?
Any model trained on human images inherits human preferences, and BecomeTen does not correct for ethnicity or gender identity; the rubric is fixed and applied to the photo as given. That is stated openly rather than hidden. The score is a position on one scale, not a claim about worth or about how any group is seen.
What is the difference between the score and the potential?
The score is the estimate of the photo now. Potential is the estimate of what is reachable without surgery, capped per area in code: harmony +6, eye area +10, jaw and chin +14, skin +25, hair +20. The gap between the two is what the plan is written to close, and it is often smaller than people expect.