Claude Code skill · labels by TypeSafe Jev · MIT
From review noise
to ranked evidence.
Analyze public product reviews and give your team an evidence-backed starting point for your product research.
- Reads reviews from an App Store, Google Play, or Steam page: 300 by default, up to 10,000.
- Labels each one with TypeSafe.ai's Jev and sorts complaints into your product's own areas.
- Ranks the areas, with verbatim quotes as evidence.
A ranked report with real quotes
Inside the sample report
- Summarywhat to fix first
- Key numberssentiment and ratings
- Product areasranked, with quotes
- Reviewsbugs, requests, churn
- Versionsnewest first
- Methodhow it was made
Each run savesreport.htmlreview_labels.csvbrief.md
Use it to check a release, size up a competitor, read a Steam patch's reception, or sort support tickets from a CSV export.
Open the sample report ↗Code counts, Jev labels, Claude explains
1Review in
CODEfetches it from the store page
2Jev labels it
JEVanswers about 45 questions per review
3Counted by area
- Sign-up and login23
- Design20
- Pricing19
- Reminders13
CODEweights by severity and ranks
Illustration only. A run ends with the HTML report shown above. Quotes and counts are real, from 300 Todoist reviews on Google Play.
CLAUDEdrafts the product areas for your app, checks the results, and writes the summary from the numbers.
Every quote is a sentence copied from a review. Near 50/50 labels are left out of the counts and listed for a person to check.
How the numbers are made · for researchers
Every Jev question, threshold, and weight is in one file, jev_questions.py, where you can check or change them.
- When a label counts
- A label counts at 60% probability or higher. Between 40% and 60% it's left out of the counts and listed as borderline. Off-topic needs 75%, because it removes a review from every count.
- Severity
- Minor: cosmetic, or a dislike of a design or business decision. Degraded: works poorly, or a feature is missing or removed. Blocking: the app or main task doesn't work, or the reviewer lost data or money.
- Priority
- Each problem review adds 1 + severity (0–3) + 2 if the reviewer mentions leaving + 1 if they tie it to an update. Areas rank by that sum.
- Likely rank
- Reviews are resampled 500 times with a fixed seed. "Likely #1–3" is the middle 90% of an area's ranks; overlapping ranges mean the order isn't settled.
- Trends
- Reviews are split at the median date. "Fewer" or "more lately" needs a one-sided Fisher exact p < 0.025, Holm-adjusted for the number of areas tested, and at least 6 complaints.
- Counts
- One review can count toward several areas, so area counts add up to more than the number of problem reviews.
What we've tested
- Repeatability. 5 reruns on the same 300 reviews for 3 apps changed 0.2–0.4% of label decisions; the top 3 areas held (one app swapped #2 and #3).
- Sampling noise. At 300 reviews only the top 2 to 4 areas were reliably ordered; at 10,000 the top 4 ranges were one or two places wide.
What we haven't tested
- Agreement with human coders. No study yet compares Jev's labels with a researcher's codes. Spot-check quotes and borderline reviews before you cite a number.
- Who writes reviews. Reviewers self-select and lean toward strong feelings. The report describes them, not your whole user base.
Set it up once, then paste a link
You need Claude Code and a TypeSafe API key (TypeSafe bills per use). You won't write code.
1Set up once4 steps · about 5 min
Paste each command into a terminal and press Return.
Clone the skill into your skills folder
git clone https://github.com/Ying8109/app-review-triage.git ~/.claude/skills/app-review-triage
If your Mac asks to install developer tools, say yes. For one project only, clone into that project's
.claude/skills/instead.Install uv
curl -LsSf https://astral.sh/uv/install.sh | sh
Or
brew install uv. On Windows or without uv, see the README.Add your TypeSafe key to
~/.zshrcor~/.bashrcexport TYPESAFE_API_KEY="your-key-here"On Windows:
setx TYPESAFE_API_KEY "your-key-here". Never paste the key into a chat.Start a new Claude Code session
From a new terminal, or a new session in the desktop app. Claude Code finds the skill on its own.
2Each runClaude checks in 3 times
Paste a review link
Ask in your own words, or type
/app-review-triage.Answer four questions
Is your key set (never the key itself), which product and link, what to categorize, and how many reviews. Claude repeats your answers back.
Approve the run plan
Dates, product areas, what goes to TypeSafe and Anthropic, and the estimated time and cost. Nothing goes to Jev until you say "run it".
Open the report
Claude opens it in your browser and recaps what you asked for against what was analyzed.
3Time and cost300 reviews ≈ 6 s, $0.10
| Reviews | Model time | Jev cost | Report |
|---|---|---|---|
| 300 (default) | ~6 s | ~$0.10 | ~0.5 MB |
| 1,000 | ~20 s | ~$0.30 | ~1.2 MB |
| 5,000 | ~1.5 min | ~$1.50 | ~5 MB |
| 10,000 (max) | ~3 min | ~$3 | ~10 MB |
Estimates from test runs with jev-1.13.0 and about 18 areas; yours vary with review length and TypeSafe's pricing. Apple's public feed stops at 500 reviews, so use an App Store Connect export for more. Interrupted runs resume where they stopped.
Questions
What leaves my machine?
Sent to
TypeSafe, for Jev's labels
- Review title
- Review text
- App name
Not sent
- Star rating
- Date
- App version
Sent to
Anthropic, while Claude works
- A sample of reviews
- The summary: counts, ratings, quotes
Your secret
TypeSafe API key
- Read from your environment
- Never printed or saved to a file
- Claude never asks for it
Reviews are treated as untrusted. Fields are escaped, formula-like CSV cells get an apostrophe, and Claude is told to ignore instructions inside reviews. This lowers the risk but can't remove it.
Before you share a report, remember it contains every analyzed review in full, which may include an email address or order number. Private exports end up in it too.
What does a run cost?
How accurate are the labels?
Can I put these numbers in a readout?
Can I use my own product areas?
Which languages work?
What about review sites like Trustpilot?
Tell me how it went
Email me what you used it for, what worked, and what confused you.
- Which app or store did you run it on, and how many reviews?
- Did the ranked areas match what your team already knew?
- Where did setup or a run get stuck?
- What would make you use it again?
Opens your email app with a short template. You can also write to yingchenpm at gmail dot com. Please don't include your API key or private review data.
For bugs, open a GitHub issue so others can see the fix.
See what your reviewers want fixed.
Free to download. MIT license.