app-review-triage GitHub

Claude Code skill · labels by TypeSafe Jev · MIT

From review noise
to ranked evidence.

Analyze public product reviews and give your team an evidence-backed starting point for your product research.

  1. Reads reviews from an App Store, Google Play, or Steam page: 300 by default, up to 10,000.
  2. Labels each one with TypeSafe.ai's Jev and sorts complaints into your product's own areas.
  3. Ranks the areas, with verbatim quotes as evidence.
~6 s · ~$0.10estimated model time and Jev cost for 300 reviews
10,000reviews per run, at most
3 review sourcesApp Store, Google Play, and Steam, plus CSV/JSON exports
▶ 1-minute intro, with sound
01 / What you get

A ranked report with real quotes

Top of the sample report: the Todoist header, a 'what to fix first' summary naming sign-up and login, design, and pricing, and key-number tiles for 300 reviews.
Sample report: 300 public Google Play reviews of Todoist, shown without reviewer names.

Inside the sample report

  1. Summarywhat to fix first
  2. Key numberssentiment and ratings
  3. Product areasranked, with quotes
  4. Reviewsbugs, requests, churn
  5. Versionsnewest first
  6. Methodhow it was made

Each run savesreport.htmlreview_labels.csvbrief.md

Use it to check a release, size up a competitor, read a Steam patch's reception, or sort support tickets from a CSV export.

Open the sample report ↗
02 / How it works

Code counts, Jev labels, Claude explains

Fig. 1 · How one review moves through the skill

1Review in

CODEfetches it from the store page

2Jev labels it

JEVanswers about 45 questions per review

3Counted by area

  1. Sign-up and login23
  2. Design20
  3. Pricing19
  4. Reminders13

CODEweights by severity and ranks

Illustration only. A run ends with the HTML report shown above. Quotes and counts are real, from 300 Todoist reviews on Google Play.

CLAUDEdrafts the product areas for your app, checks the results, and writes the summary from the numbers.

Every quote is a sentence copied from a review. Near 50/50 labels are left out of the counts and listed for a person to check.

How the numbers are made · for researchers

Every Jev question, threshold, and weight is in one file, jev_questions.py, where you can check or change them.

When a label counts
A label counts at 60% probability or higher. Between 40% and 60% it's left out of the counts and listed as borderline. Off-topic needs 75%, because it removes a review from every count.
Severity
Minor: cosmetic, or a dislike of a design or business decision. Degraded: works poorly, or a feature is missing or removed. Blocking: the app or main task doesn't work, or the reviewer lost data or money.
Priority
Each problem review adds 1 + severity (0–3) + 2 if the reviewer mentions leaving + 1 if they tie it to an update. Areas rank by that sum.
Likely rank
Reviews are resampled 500 times with a fixed seed. "Likely #1–3" is the middle 90% of an area's ranks; overlapping ranges mean the order isn't settled.
Trends
Reviews are split at the median date. "Fewer" or "more lately" needs a one-sided Fisher exact p < 0.025, Holm-adjusted for the number of areas tested, and at least 6 complaints.
Counts
One review can count toward several areas, so area counts add up to more than the number of problem reviews.

What we've tested

  • Repeatability. 5 reruns on the same 300 reviews for 3 apps changed 0.2–0.4% of label decisions; the top 3 areas held (one app swapped #2 and #3).
  • Sampling noise. At 300 reviews only the top 2 to 4 areas were reliably ordered; at 10,000 the top 4 ranges were one or two places wide.

What we haven't tested

  • Agreement with human coders. No study yet compares Jev's labels with a researcher's codes. Spot-check quotes and borderline reviews before you cite a number.
  • Who writes reviews. Reviewers self-select and lean toward strong feelings. The report describes them, not your whole user base.
03 / Get started

Set it up once, then paste a link

You need Claude Code and a TypeSafe API key (TypeSafe bills per use). You won't write code.

1Set up once4 steps · about 5 min

Paste each command into a terminal and press Return.

  1. Clone the skill into your skills folder

    git clone https://github.com/Ying8109/app-review-triage.git ~/.claude/skills/app-review-triage

    If your Mac asks to install developer tools, say yes. For one project only, clone into that project's .claude/skills/ instead.

  2. Install uv

    curl -LsSf https://astral.sh/uv/install.sh | sh

    Or brew install uv. On Windows or without uv, see the README.

  3. Add your TypeSafe key to ~/.zshrc or ~/.bashrc

    export TYPESAFE_API_KEY="your-key-here"

    On Windows: setx TYPESAFE_API_KEY "your-key-here". Never paste the key into a chat.

  4. Start a new Claude Code session

    From a new terminal, or a new session in the desktop app. Claude Code finds the skill on its own.

2Each runClaude checks in 3 times
  1. Paste a review link

    Ask in your own words, or type /app-review-triage.

  2. Answer four questions

    Is your key set (never the key itself), which product and link, what to categorize, and how many reviews. Claude repeats your answers back.

  3. Approve the run plan

    Dates, product areas, what goes to TypeSafe and Anthropic, and the estimated time and cost. Nothing goes to Jev until you say "run it".

  4. Open the report

    Claude opens it in your browser and recaps what you asked for against what was analyzed.

3Time and cost300 reviews ≈ 6 s, $0.10
ReviewsModel timeJev costReport
300 (default)~6 s~$0.10~0.5 MB
1,000~20 s~$0.30~1.2 MB
5,000~1.5 min~$1.50~5 MB
10,000 (max)~3 min~$3~10 MB

Estimates from test runs with jev-1.13.0 and about 18 areas; yours vary with review length and TypeSafe's pricing. Apple's public feed stops at 500 reviews, so use an App Store Connect export for more. Interrupted runs resume where they stopped.

04 / FAQ

Questions

What leaves my machine?

Sent to

TypeSafe, for Jev's labels

  • Review title
  • Review text
  • App name

Not sent

  • Star rating
  • Date
  • App version

Sent to

Anthropic, while Claude works

  • A sample of reviews
  • The summary: counts, ratings, quotes

Your secret

TypeSafe API key

  • Read from your environment
  • Never printed or saved to a file
  • Claude never asks for it

Reviews are treated as untrusted. Fields are escaped, formula-like CSV cells get an apostrophe, and Claude is told to ignore instructions inside reviews. This lowers the risk but can't remove it.

Before you share a report, remember it contains every analyzed review in full, which may include an email address or order number. Private exports end up in it too.

What does a run cost?
In our tests, Jev cost roughly $0.06 to $0.10 per 300 reviews. Treat that as an estimate. Longer reviews and more product areas cost more, and TypeSafe sets its own prices. Claude's own steps count toward your usual Claude Code usage.
How accurate are the labels?
Repeat runs agree closely, but no study yet compares the labels with human coders. Near 50/50 labels are left out of the counts and listed for a person to check, and each rank comes with a likely range. See How the numbers are made.
Can I put these numbers in a readout?
Yes, as a triage of what reviewers say. Quote the rank range with each rank, check the quotes you use against the review, and say the sample is public reviewers, not all users.
Can I use my own product areas?
Yes. Name your features or areas and the skill tells Claude to use them as you wrote them, adding only what's missing. You can also shape the areas around one research question.
Which languages work?
English works best. Other languages are labeled but flagged in the report.
What about review sites like Trustpilot?
Pages with structured review data often work directly. For others, Claude can try reading the page in a browser, but some sites block this or show only a few hundred reviews.
05 / Feedback

Tell me how it went

Email me what you used it for, what worked, and what confused you.

  • Which app or store did you run it on, and how many reviews?
  • Did the ranked areas match what your team already knew?
  • Where did setup or a run get stuck?
  • What would make you use it again?
Email your feedback

Opens your email app with a short template. You can also write to yingchenpm at gmail dot com. Please don't include your API key or private review data.

For bugs, open a GitHub issue so others can see the fix.

See what your reviewers want fixed.

Free to download. MIT license.