Methodology
How Appeye turns public app reviews into every metric on the site.
How fresh is this?
An analysis window is the exact set of quality-filtered reviews behind an app's score: how many reviews were analysed, the first and last review dates, and how many distinct days contain reviews. It is evidence coverage, not the date the app record was written.
Scoring runs continuously, and every app page carries its own scored date. In the current published catalogue, 25,699 apps have an App Store analysis window, 39,583 have a Google Play window, and 3,963 have a combined window. A scored date is present for 61,304 of 61,305 apps.
Where the data comes from
Two public sources: Apple's App Store and Google Play. Appeye collects app metadata (name, publisher, category, rating, rating count) and the text of user reviews. Nothing is bought, inferred, or supplied by app developers.
Which apps appear
An app is published on Appeye once it has at least 100 usable written reviews. Below that, there isn't enough text to say anything meaningful, so the app stays out of the catalogue rather than appearing with a thin score.
Which reviews count
Not every review is usable. Before anything is calculated, a review is excluded if it:
- is shorter than 15 characters, or fewer than 4 words
- fails a gibberish check (the ratio of vowels to letters falls outside 10%–75%, which catches keyboard mashing)
- is the 4th or later review from the same reviewer handle on the same app
- is not in one of the six languages Appeye scores
Appeye scores reviews in English, German, Spanish, French, Italian and Dutch. English is scored with VADER. The other five are scored by nlptown/bert-base-multilingual-uncased-sentiment, run locally on CPU.
The boundary is not an arbitrary list: those five are the languages that model was fine-tuned on for product reviews specifically. A review outside the six is collected but not scored, and does not count toward the 100-review floor — a sentiment number from a model never tuned for that language would be a figure nobody can stand behind, so Appeye does not produce one.
Non-English scoring is not a fringe case. 220,096 reviews in the published catalogue were scored by the multilingual model, and they appear in the analysis of 22,582 of 61,305 published apps — 37% — with 178 apps that would fall below the 100-review floor without them. Some rest on them almost entirely: Midatacrédito is 127 of 127 analysed reviews, and SUBE is 402 of 411 analysed reviews. Each app page lists the languages behind its own analysis.
The split is worth stating plainly, because it is small. Of everything Appeye scores in the published catalogue, 34,073,638 reviews are English and 220,096 are not — 0.6%. Within that non-English share, French and Spanish alone account for 76.2%; the three languages added most recently contribute 52,327 reviews between them, 23.8%.
A second, wider measurement covers the whole review archive, including apps that are not published — a different population from the figures above, so the two are never divided into each other. Across that archive, 36,589,051 reviews are English and 247,364 are not — 0.7% of everything scored, against 23,344,694 held back unscored, about 94 times as many. Most of that pile is language Appeye cannot score, not junk. Measured 1 October 2026.
English sits off this axis on purpose: 34,073,638 reviews, against 220,096 in the five languages above.
The rest waits rather than being guessed at: English is scored with VADER and five languages with a multilingual model fine-tuned on product reviews. Everything outside those six is held back, because a sentiment number from a model never tuned for the language is a figure nobody can stand behind.
The review count shown on an app page is the number that passed all of these — not the app's total review count on the store.
The review window
Appeye analyses the most recent reviews it has collected, up to 2,000 per store.
"Days with reviews" counts how many separate calendar dates those reviews fall on. It is not the gap between the first and the last: an app can have reviews from two dates a month apart and still only have two days of coverage.
Busy apps have shorter windows. An app receiving thousands of reviews a day fills its 2,000 with a day or two of activity, while a quieter app's 2,000 can span years. Neither store lets us request reviews by date, so the window is a consequence of how fast an app is reviewed — not a choice.
A short window is not a warning. It means the analysis reflects a recent period rather than a long history, which for a fast-moving app is often the more useful read. Every app page states its window so you know which one you are looking at.
Sentiment
Each usable review is scored for sentiment. English reviews use VADER, a rule-based analyser built for short social text, which produces a compound score from -1 to +1:
- positive: 0.05 or above
- negative: -0.05 or below
- neutral: between the two
VADER reads punctuation, capitals and negation ("not good" is not scored as "good"). It is rule-based, not a language model — fast, free and deterministic, so the same review always scores the same way.
German, Spanish, French, Italian and Dutch reviews are scored by nlptown/bert-base-multilingual-uncased-sentiment, run locally on CPU, and mapped onto the same three labels. Every scored row records which model produced it, so the split between VADER and the multilingual model is answerable per app.
Reviews in any other language are collected but never scored, and never counted toward the review totals behind a score.
How the score is built (0–100)
50% store rating, 50% review sentiment. A fixed weighting, applied the same way to every app — not a model, not tuned, not learned from anything. There are no other inputs: review volume, price, downloads, recency and category all stay out of it.
score = (store rating as a percentage + average sentiment as a percentage) / 2
A 5-star rating becomes 100; a sentiment average of +1 becomes 100. So an app rated 4.2 with mildly positive review text scores differently from one rated 4.2 whose reviews read badly — which is the point. A star rating alone hides that gap.
Because the weighting never changes, two scores are always comparable, and any difference between them comes from the two inputs rather than from how the number was calculated.
If either half is missing, there is no score. An app with a rating but no usable review text gets nothing rather than a rating-only number. This is deliberate: a partial score is not a smaller claim, it is a different one, and a score built from ratings alone once ranked metadata-only apps at the top.
How much data each number rests on
A score of 84 built on 118 reviews and a score of 84 built on 1,959 are not the same claim, so wherever a score appears Appeye states the number of analysed reviews behind it, and glosses that count in words:
No app is published on fewer than 100 analysed reviews. That floor is a rule, not a result: an app whose App Store and Google Play analysed review counts do not add up to 100 stays out of the catalogue entirely, so there is no band of thinly-evidenced apps to report. Above the floor, the gloss reads:
- Moderate — 100 to 399 (30,149 apps, 49%)
- High — 400 or more (31,156 apps, 51%)
The count is the App Store and Google Play analysed review counts added together, and the threshold between the two words is measured across the 61,305 published apps rather than chosen for effect: 400 is four times the publishing floor and sits just under the 75th percentile, which is 808.
Every figure in this section is counted from the catalogue when this page loads, not written into the text.
There is no confidence percentage, and there never will be. "96% confident" would imply a statistical interval nobody has computed. These three words are a reading of a count you can see, nothing more — the count is the evidence, and it is always shown.
Sample size
Every sentiment figure on the site names the number of substantive reviews it was computed from. A percentage without its sample size is not a finding, so the two are never separated: 62% negative from 89 reviews is reported that way, not as 62%.
Those counts are of quality-filtered reviews only. Before anything is calculated a review has to clear a minimum length, pass a gibberish check, not be a repeat from the same reviewer handle, and be in a language Appeye scores — the full rules are under which reviews count. So the sample size is always smaller than the store's headline review total, and it is the smaller number Appeye reports.
Sample sizes are also per store. An app can have thousands of Google Play reviews analysed and a few dozen on the App Store, which is why per-store figures carry their own counts instead of borrowing the combined one.
Themes
Themes are the two-word phrases that recur in reviews, counted separately for positive and negative ones, with common filler words removed. Up to 8 phrases per group are shown, and a phrase must appear at least twice to qualify.
They are counted, not interpreted. "battery life" appearing 12 times in negative reviews means 12 people wrote those two words near each other — it does not mean Appeye has judged the battery to be bad.
Churn signal
An estimate of how many reviewers sound like they are leaving. It combines:
- the percentage of reviews that are negative
- how often reviews contain departure language — around 30 tracked phrases such as "uninstalled", "switching to", "cancelled my", "want my money back", "found a better app" (each hit adds 2 points, capped at 20)
- the direction of travel: comparing the negative percentage over the last 3 months against the 3 months before, a shift of 5 percentage points or more counts as worsening (+15) or improving (-10)
The result is capped to 0–100 and reported as HIGH (50+), MEDIUM (25–49) or LOW (under 25), together with the phrases that triggered it, so you can see what it is reacting to.
Trend needs at least 6 months of review history. With less, it reports "not enough history" rather than guessing.
Reviewer breakdown
The distribution of 1- to 5-star ratings across the reviews analysed. This is computed from the reviews Appeye holds, not from the store's own histogram.
User engagement
How often a developer replies to reviews, and how quickly.
This is Google Play only, and always will be. Apple does not expose developer replies to third parties, so there is nothing to measure on the App Store side. An empty engagement panel on an App Store card is a limitation of the data source, not a missing reply from the developer.
Category rank
Within each category, apps are ranked by score, highest first, separately for each store. The rank shown is the app's position among apps in the same category with a score.
Recommendations
Two kinds, and they coexist permanently:
- Rule-based (free, shown by default): generated from the app's own theme and churn signals by fixed rules. No AI is involved — this was a deliberate engineering choice so the feature could stay free.
- AI reports (paid, on request): a written analysis generated by a language model over the same data.
Neither replaces the other.
What Appeye does not do
- It does not score reviews outside English, German, Spanish, French, Italian and Dutch — the six the analysers are built or fine-tuned for. Reviews in other languages are collected but not scored, so they do not affect any number on the site.
- It does not show an App Store star distribution. Apple's public API returns only an average rating and a total count, never a per-star breakdown. Google Play does return one, which is why the two cards differ.
- It has no "Family" category. Apple has no top-level Family chart, and Google Play reports Family apps under Educational or Education, so the category cannot be built from either source.
- It does not contact developers, accept paid placement, or adjust any score on request.
When Appeye shows nothing
Where there isn't enough data, Appeye shows nothing rather than a number. An empty field means the calculation could not be made honestly — not that the app scored zero.
Want to challenge a number or ask about the method? Contact Appeye.