How we score AI meeting assistants

This page explains concretely what makes a good AI meeting assistant, which metrics we use to measure the quality of its summaries, and how documented data turns into a reproducible score from 0 to 5. One rubric, the same for every assistant.

Our guiding principle: the result follows from the data, never the data from a desired result.

Every AI meeting assistant runs through the same public rubric. No bonus points, no downgraded competitors, and every hard claim carries a source and a check date.

What makes a good summary

Before ticking off features we settle the important question: when is the output actually good? We anchor it to six measurable properties.

Transcription accuracy (WER)

If the transcript is wrong, everything built on it is wrong. The Word Error Rate (WER) is the share of words the AI assistant gets wrong.

WER = (substitutions + deletions + insertions) / words spoken. Lower is better.

Strongunder 10 %Usable10 to 20 %Weakover 20 %

Speaker separation (DER)

A summary is only useful if it knows who said and promised what. The Diarization Error Rate (DER) is the share of talk time attributed to the wrong speaker.

Share of talk time labelled with the wrong speaker. Matters most in group calls.

Strongunder 10 %Usable10 to 20 %Weakover 20 %

Faithfulness (no hallucinations)

The summary may only contain what was actually said. Faithfulness is the share of statements in the summary that are backed by the transcript. The opposite is the hallucination rate.

Backed statements / all statements. We reward AI assistants that link each point back to the exact spot in the transcript.

Strongover 95 % backedUsable85 to 95 %Weakunder 85 %

Tasks & decisions (recall / precision)

The real payoff of a meeting assistant is catching every to-do and decision without inventing any. Recall = how many real tasks it caught. Precision = how many caught tasks were real.

We check the extracted action items against the actual conversation. High on both, not just one.

Strongrecall & precision highUsableone of the two weakWeakmisses or invents tasks

Signal density

A good summary is short enough to skim yet covers every decision, task and open question. Too long is as unhelpful as too short.

Coverage of the key points relative to length. Rambling recaps and one-liners both lose points.

Strongconcise, completeUsabletoo long or thinWeakmisses key points

Time to result

A summary that lands the next day is a summary nobody reads. It should be ready right after the call, ideally within minutes.

How long after the meeting the finished summary is available.

StrongminutesUsablewithin the hourWeakhours or a day

These metrics are the yardstick. We cannot re-measure them for every AI assistant in every language ourselves, so we translate them into a high, medium or low rating drawn from vendor benchmarks, independent tests and our own spot checks, including in German. We label such values as indicative.

The seven scoring dimensions

Every AI meeting assistant is scored across seven dimensions. The percentages are the default weight: how strongly each dimension feeds the overall grade.

Security & compliance25 %

The most heavily weighted block for regulated and privacy-conscious buyers.

Capture & transcription18 %

The core job: turning a conversation into an accurate, readable transcript.

AI insights18 %

What the AI assistant makes of the conversation: summaries, tasks and searchable knowledge.

Integrations & workflow15 %

How well the AI assistant plugs into the systems a team already uses.

Usability & onboarding9 %

How quickly a non-technical team gets value out of the AI assistant.

Pricing & model8 %

Cost and transparency. A free plan matters to individuals; serious teams weigh value and clarity.

Vendor & trust7 %

Who stands behind the AI assistant and how reachable they are.

The default weights reflect our audience: privacy-conscious mid-market companies in the German-speaking region. That is why security & compliance carries the most weight. If your priorities differ, the buyer profiles further down re-weight everything.

How data becomes a grade

The overall grade is not handed out, it is computed. In three steps, reproducible from published rules and documented data.

1. Every value onto 0 to 1

Yes = 1, partial = 0.5, no = 0. Enum values on a fixed, published scale. Numbers on a defined range (languages normalised, price inverted: cheaper is better). Not confirmed does not count and lowers data completeness instead.

2. Average per dimension

The normalised values of a dimension form its dimension score from 0 to 1, shown as a bar from 0 to 100.

3. Weighted sum into the grade

Each dimension score is multiplied by its weight, summed, and mapped onto the familiar 0 to 5 scale.

A worked example

A made-up example assistant that shows the maths without making a claim about a real product. Weight times dimension score, summed and scaled to 0 to 5.

DimensionWeightScore (0–1)Contribution
Security & compliance25 %0.800.20
Capture & transcription18 %0.900.16
AI insights18 %0.750.14
Integrations & workflow15 %0.700.10
Usability & onboarding9 %0.850.08
Pricing & model8 %0.600.05
Vendor & trust7 %0.650.05
Sum (0–1) × 50.77 × 5 = 3.86

This replaces the earlier hand-set rating. Nobody can accuse us of assigning grades freely, because they follow from the rubric.

The scoring scales in detail

For multi-level criteria we publish the scale. This is the exact mapping we compute with.

Data residency

  • EU (mandatory)1.00
  • EU (optional)0.70
  • Global0.40
  • USA0.20

Customer data used for AI training

  • Never1.00
  • Opt-out0.60
  • Opt-in required0.80
  • Yes0.00

Legal jurisdiction

  • EU1.00
  • UK0.70
  • USA0.40
  • Other0.40

SOC 2

  • Type II1.00
  • Type I0.50
  • Assessed compliant0.50
  • No0.00

Transcription accuracy

  • High1.00
  • Medium0.60
  • Low0.30

Price transparency

  • Public1.00
  • Partial0.50
  • On request0.20

Buyer profiles: same data, different weights

Not everyone weights the same. So we offer switchable profiles. They change no rubric and no raw data, only the percentages. Everyone can see why the order shifts.

Privacy first

Security & compliance pushed to roughly 40 %. For regulated teams and works councils.

Sales / CRM

Integrations and CRM depth weighted up, for teams that log every call in their CRM.

Budget conscious

Price and a permanent free plan weighted up, for individuals and small teams.

Precision / minutes

Capture and AI insights weighted up, when the transcript itself is the deliverable.

Evidence, sources and updates

A rubric is only as good as the data inside it. So we keep the origin of every figure verifiable.

Data sources

Public vendor documentation, pricing pages, help centres, certificate registries, DPA documents and our own spot checks. Fast-changing figures such as price or language count are labelled indicative.

Evidence for hard claims

Certifications, hosting location, data use and price transparency carry a clickable source and a check date, visible right in the comparison table.

Data completeness shown

Missing values are marked as not confirmed rather than guessed. That lowers completeness but never quietly drags an AI assistant down or looks like a confirmed no.

Regular review

We re-check the data in cycles and show a last-reviewed date (last on 2026-06-01). Vendor offerings change fast, so confirm details on the vendor site before buying.

Operator disclosure

MeetingSummary is operated by Aliru GmbH, which also offers the meeting assistant Sally AI. That is exactly why this rubric is public and identical for every AI assistant. Sally AI is scored on the same criteria and weights as any other assistant, with no bonus points and no downgraded competitors. Where Sally AI ranks high under a profile, it is because of documented strengths such as EU hosting, GDPR and support in the German-speaking region, not a thumb on the scale.