How we score AI meeting assistants
This page explains concretely what makes a good AI meeting assistant, which metrics we use to measure the quality of its summaries, and how documented data turns into a reproducible score from 0 to 5. One rubric, the same for every assistant.
Our guiding principle: the result follows from the data, never the data from a desired result.
Every AI meeting assistant runs through the same public rubric. No bonus points, no downgraded competitors, and every hard claim carries a source and a check date.
What makes a good summary
Before ticking off features we settle the important question: when is the output actually good? We anchor it to six measurable properties.
Transcription accuracy (WER)
If the transcript is wrong, everything built on it is wrong. The Word Error Rate (WER) is the share of words the AI assistant gets wrong.
WER = (substitutions + deletions + insertions) / words spoken. Lower is better.
Speaker separation (DER)
A summary is only useful if it knows who said and promised what. The Diarization Error Rate (DER) is the share of talk time attributed to the wrong speaker.
Share of talk time labelled with the wrong speaker. Matters most in group calls.
Faithfulness (no hallucinations)
The summary may only contain what was actually said. Faithfulness is the share of statements in the summary that are backed by the transcript. The opposite is the hallucination rate.
Backed statements / all statements. We reward AI assistants that link each point back to the exact spot in the transcript.
Tasks & decisions (recall / precision)
The real payoff of a meeting assistant is catching every to-do and decision without inventing any. Recall = how many real tasks it caught. Precision = how many caught tasks were real.
We check the extracted action items against the actual conversation. High on both, not just one.
Signal density
A good summary is short enough to skim yet covers every decision, task and open question. Too long is as unhelpful as too short.
Coverage of the key points relative to length. Rambling recaps and one-liners both lose points.
Time to result
A summary that lands the next day is a summary nobody reads. It should be ready right after the call, ideally within minutes.
How long after the meeting the finished summary is available.
These metrics are the yardstick. We cannot re-measure them for every AI assistant in every language ourselves, so we translate them into a high, medium or low rating drawn from vendor benchmarks, independent tests and our own spot checks, including in German. We label such values as indicative.
The seven scoring dimensions
Every AI meeting assistant is scored across seven dimensions. The percentages are the default weight: how strongly each dimension feeds the overall grade.
The most heavily weighted block for regulated and privacy-conscious buyers.
The core job: turning a conversation into an accurate, readable transcript.
What the AI assistant makes of the conversation: summaries, tasks and searchable knowledge.
How well the AI assistant plugs into the systems a team already uses.
How quickly a non-technical team gets value out of the AI assistant.
Cost and transparency. A free plan matters to individuals; serious teams weigh value and clarity.
Who stands behind the AI assistant and how reachable they are.
The default weights reflect our audience: privacy-conscious mid-market companies in the German-speaking region. That is why security & compliance carries the most weight. If your priorities differ, the buyer profiles further down re-weight everything.
How data becomes a grade
The overall grade is not handed out, it is computed. In three steps, reproducible from published rules and documented data.
1. Every value onto 0 to 1
Yes = 1, partial = 0.5, no = 0. Enum values on a fixed, published scale. Numbers on a defined range (languages normalised, price inverted: cheaper is better). Not confirmed does not count and lowers data completeness instead.
2. Average per dimension
The normalised values of a dimension form its dimension score from 0 to 1, shown as a bar from 0 to 100.
3. Weighted sum into the grade
Each dimension score is multiplied by its weight, summed, and mapped onto the familiar 0 to 5 scale.
A worked example
A made-up example assistant that shows the maths without making a claim about a real product. Weight times dimension score, summed and scaled to 0 to 5.
| Dimension | Weight | Score (0–1) | Contribution |
|---|---|---|---|
| Security & compliance | 25 % | 0.80 | 0.20 |
| Capture & transcription | 18 % | 0.90 | 0.16 |
| AI insights | 18 % | 0.75 | 0.14 |
| Integrations & workflow | 15 % | 0.70 | 0.10 |
| Usability & onboarding | 9 % | 0.85 | 0.08 |
| Pricing & model | 8 % | 0.60 | 0.05 |
| Vendor & trust | 7 % | 0.65 | 0.05 |
| Sum (0–1) × 5 | 0.77 × 5 = 3.86 | ||
This replaces the earlier hand-set rating. Nobody can accuse us of assigning grades freely, because they follow from the rubric.
The scoring scales in detail
For multi-level criteria we publish the scale. This is the exact mapping we compute with.
Data residency
- EU (mandatory)1.00
- EU (optional)0.70
- Global0.40
- USA0.20
Customer data used for AI training
- Never1.00
- Opt-out0.60
- Opt-in required0.80
- Yes0.00
Legal jurisdiction
- EU1.00
- UK0.70
- USA0.40
- Other0.40
SOC 2
- Type II1.00
- Type I0.50
- Assessed compliant0.50
- No0.00
Transcription accuracy
- High1.00
- Medium0.60
- Low0.30
Price transparency
- Public1.00
- Partial0.50
- On request0.20
Buyer profiles: same data, different weights
Not everyone weights the same. So we offer switchable profiles. They change no rubric and no raw data, only the percentages. Everyone can see why the order shifts.
Privacy first
Security & compliance pushed to roughly 40 %. For regulated teams and works councils.
Sales / CRM
Integrations and CRM depth weighted up, for teams that log every call in their CRM.
Budget conscious
Price and a permanent free plan weighted up, for individuals and small teams.
Precision / minutes
Capture and AI insights weighted up, when the transcript itself is the deliverable.
Evidence, sources and updates
A rubric is only as good as the data inside it. So we keep the origin of every figure verifiable.
Data sources
Public vendor documentation, pricing pages, help centres, certificate registries, DPA documents and our own spot checks. Fast-changing figures such as price or language count are labelled indicative.
Evidence for hard claims
Certifications, hosting location, data use and price transparency carry a clickable source and a check date, visible right in the comparison table.
Data completeness shown
Missing values are marked as not confirmed rather than guessed. That lowers completeness but never quietly drags an AI assistant down or looks like a confirmed no.
Regular review
We re-check the data in cycles and show a last-reviewed date (last on 2026-06-01). Vendor offerings change fast, so confirm details on the vendor site before buying.
Operator disclosure
MeetingSummary is operated by Aliru GmbH, which also offers the meeting assistant Sally AI. That is exactly why this rubric is public and identical for every AI assistant. Sally AI is scored on the same criteria and weights as any other assistant, with no bonus points and no downgraded competitors. Where Sally AI ranks high under a profile, it is because of documented strengths such as EU hosting, GDPR and support in the German-speaking region, not a thumb on the scale.