Home / Why sentiment misses it
The core argument
Why sentiment analysis
misses your worst calls
Because the worst calls are handled beautifully. A capable agent stays warm and professional while your business quietly fails the person on the other end. Tone scoring hears the warmth. It cannot hear the failure.
The measurement
We checked this properly, on real calls
Over 30 days at an Australian corporate travel agency, WiseSentry reviewed every recorded call in scope and raised 45 to management. We then looked at what the phone system's own sentiment score had said about those same 45 calls.
On a 1–100 scale where higher is friendlier, the calls that genuinely needed a manager sat at a median of 82. Only three of the forty-five fell into the range a "low sentiment" alert would catch.
Put the other way round: if this business had been monitoring sentiment and acting on scores below 40, it would have seen 3 of its 45 real problems — and would have been reassured by the other 42.
Why this happens
Good service and a bad outcome sound identical
The agent is doing their job
They apologise, they stay calm, they promise to chase it. Every acoustic and lexical signal a sentiment model looks for says "this is going fine".
Australians under-complain
"No worries", "that's alright", "no problem at all" — said by someone who will quietly move their business elsewhere next quarter.
The failure is in the facts
A promised callback that never came is a fact about the past, not a tone in the present. No amount of listening to how it was said reveals it.
And the inverse is just as expensive
Sentiment doesn't only miss real problems — it manufactures false ones.
A customer swearing cheerfully about an airline strike, to an agent who is fixing it brilliantly, will score badly. Chase that alert and you have spent a manager's afternoon on a call where everyone did well. Do that a few times and the alerts get filtered to a folder — which is how monitoring tools die.
Comparison
Where each approach actually looks
| Signal | What it can see | What it cannot see |
|---|---|---|
| Sentiment score | Tone, profanity, emotional intensity | A calm conversation in which the business failed |
| Keyword spotting | Named words — "cancel", "complaint", "manager" | The same meaning expressed in any other words |
| QA sampling | Deep detail — on the 1–2% of calls reviewed | The other 98%, which is where the problem usually is |
| WiseSentry | What was actually said, judged against whether your business let the customer down | Anything not said on a recorded call |
WiseSentry reads the transcript of the conversation. It does not analyse voice acoustics, stress or cadence — the signal it uses is meaning, not tone.
What we found instead
One theme, not forty-five scattered problems
When you review on content rather than tone, the flags cluster — and the cluster is the actionable finding.
In that month, 36 of the 45 flags were the same category of failure: service breakdowns, overwhelmingly traced back to one internal system that kept letting customers down in the same way. Not forty-five bad calls. One problem, surfacing forty-five times.
That is the difference between a monitoring tool and a management tool. A score tells you a call felt bad. A pattern tells you what to go and fix on Monday.
Being straight with you
What we have not measured
We can tell you what WiseSentry found. We cannot yet tell you what it missed.
Precision — whether the calls it raises are genuinely worth raising — has been checked against a client's own verdicts, call by call. Recall — the share of all real problems that it catches — has not been measured, because doing so honestly means a human listening to a large random sample of calls that were not flagged.
Any vendor quoting you a recall figure without describing that exercise is guessing. We would rather say so.
Run it over your own calls
The fastest way to test any of this is on your own recordings. We will show you what it finds and what it ignored.