Field notes · Method

Composite scores are not averages wearing a suit

Handwritten notes beside a coffee cup

When a team says they “already have an engagement score,” they usually mean this: take five metrics, min-max them, add them, maybe multiply by a recency decay, and publish a number between 0 and 100. It looks like Engagement Score Analytics. It behaves like an average that learned to wear a suit.

The tell is simple. Two users land on 62. One paid a bill yesterday and has not opened a marketing push in months. The other opens every push, checks the balance twice a day, and has not funded the wallet since November. The composite cannot tell them apart, so the lifecycle tool treats them as the same chair.

What the sum is hiding

Normalisation pretends units are comparable. A balance check and a bill payment are not comparable, even after you squeeze both onto a 0–1 scale. You have only made the arithmetic legal. You have not made the meaning honest.

In Resonance Ledger we ask chairs to pick one user story that would embarrass the current score. Almost every wallet team finds a “ghost opener”: high frequency, no economic event. If that person can outscore a quiet payer, the composite is a usage trophy, not an engagement measure.

A quieter construction

We do not ban composites. We ban unexamined ones. The Atlas version looks like this:

  1. Write the definition first. “Engagement means a user still uses the wallet to move money in ways we can observe.” That sentence excludes push opens unless you can argue they predict movement.
  2. Split signals into economic, navigational, and campaign-induced. Campaign-induced events may enter a separate diagnostic, not the score that triggers retention spend.
  3. Weight inside each family, then combine families with documented ratios. The ratio is a product decision. Write the date and the owner.
  4. Keep a residual: a flag for users whose family scores disagree. Those people are not “average.” They are the ones your band will mis-fire on.

None of this requires a new warehouse. It requires a memo. Most teams skip the memo because the dashboard already drew a sparkline.

The meeting test

If a country manager cannot retell why 62 is not one kind of person, the score is not ready. In Bangkok we run that retell out loud. It is slower than adding columns. It is also why Cohort 11 retired push opens from a wallet score that had survived two years of internal applause.

Averages are fine for temperature. Engagement is closer to a diagnosis. Treat the composite as a claim you must defend, not as a summary that arrived fully dressed.

See how Resonance Ledger practices this · More field notes