How we count

Every ranking on HiveTried comes from one of two sources: public Reddit threads or shopper reviews at Ulta and Sephora. So far that's 5,246 first-hand reports from 1,433 threads and 4,594 counted store reviews. Each page says which source it uses.

Last updated

Pages built from Reddit threads

1. Find the threads

We start from public Reddit discussions in makeup communities: r/MakeUpAddictionUK, r/Makeup, r/MakeupAddiction, r/MakeupAddictionCanada, r/Sephora, r/beauty, r/drugstoreMUA, r/muacjdiscussion. We use an archive of public Reddit posts maintained for research, so we can read whole threads without scraping Reddit itself.

For each problem we take two kinds of threads posted since January 2022:

2. Spot the products

We match product names and common nicknames ("Sky High", "Lash Princess", "Double Wear") in every comment. The 10 products mentioned most in a problem's threads get labeled.

Popular products are mentioned in hundreds of comments, so we label a random sample of up to 80 comments per product. When a product still has fewer than about 30 buyers saying whether it solved the problem, we label a second sample, this time from comments that talk about the problem itself (in either direction, "smudges" as much as "never smudges").

3. Label each report

An AI model (Claude, by Anthropic) reads each comment and answers four questions for every product it mentions:

  1. Does the comment really mean this product? ("Thrive" is also a verb.)
  2. Did the commenter use it themselves? "I've worn it for two years" counts. "I heard it's good" and "try X" do not.
  3. Did they say whether it solved the problem, and which way?
  4. Which points from a fixed list did they mention, good or bad?

The fixed lists are the same for every product of a type, so "Clumpy" means the same thing on every mascara page. We then review a sample of labels by hand before a page goes live. In our last review, we checked 40 randomly drawn labels by hand and agreed with 39 of them. The one miss counted a dry-skinned buyer's comment toward the oily-skin ranking.

4. Count people, not comments

One person counts once per product on a page, however many times they comment. Usernames are replaced by a one-way code before counting and are never stored or shown.

Pages built from store reviews

Store reviews come from people who bought the product at Ulta or Sephora, so they reach far more buyers than Reddit does. For these pages:

  1. Pick the products. We take the products with the most reviews in the category at both stores, keep the ones that are still sold, and merge the same product listed twice (shades, sizes, store listings).
  2. Search for the problem. On each product page we search the reviews for the problem ("clump", "oxidize", "waterline") and read the newest matching reviews, up to 40 per product and store. They can praise or complain: "doesn't clump at all" counts as much as "clumps".
  3. Drop paid reviews. Reviews written in exchange for a free product or a sweepstakes entry are dropped, whether the store flags them or the reviewer says so in the text.
  4. Set brand sites aside. Ulta also shows reviews first posted on brands' own websites. They run much more positive: across 19 products with both, 76% of brand-site reviewers said the product solved the problem, against 55% of Ulta shoppers. We show those numbers on the product cards, marked "not counted", and leave them out of the ranking.
  5. Label and count. The same AI model labels each review with the same questions as above, and each review counts once. In our last check of store-review labels, we checked 30 randomly drawn store-review labels against the review text and agreed with 29. The one we'd call differently: a reviewer said a foundation oxidized "a little" but liked the result, and the model counted it as oxidizing.

Product cards on these pages also show where the numbers come from, store by store, and for Sephora reviewers, by the skin type they report.

How products are ranked

A product's position comes from the share of buyers who said it solved the problem. Raw shares reward luck, though: 4 out of 4 is 100%, but tells you less than 80 out of 88. So we rank by the lower bound of a 90% confidence interval (the Wilson score), which only goes high when the share is high and enough people said so.

ProductSaid it workedShareRanking score
A4 of 4100%60
B80 of 8891%85
C30 of 5555%44

A product needs at least 8 people stating an outcome to be ranked at all.

What the percentages under each product mean

"Lengthens lashes 25%" means a quarter of that product's first-hand reports (or counted store reviews) on the page mention long lashes. Most people only mention one or two things, so these numbers are lower than satisfaction rates. Use them to compare products, and to see what people bring up most.

What this method can't tell you

We never publish anyone's comment or review, only counts and links to the public threads and product pages they came from. Brands can't pay for placement. We earn a commission from Amazon when you buy through our buttons, and that never feeds into the ranking.