dvar

Case study · sorting customer messages

Banking77

A small model sorts bank customer messages more accurately than Sonnet, at 64× lower cost.

Hinglish accuracy

59.7+2888.1

Stock Kev to two LoRA epochs. Sonnet sits at 81.2.

Cost per 1k

$0.9764×$0.015

Trained Kev on an India L4 versus billed Sonnet 4.5.

Latency

2.0s31×64ms

p50. Choice prefill versus a prompted decode.

Kev 0.5BAccuracy on the official 77 labels
60708090Sonnet 4.5 · 81.381.559.7Stockn=3,08092.788.12 LoRA epochsn=3,080
  • English
  • Hinglish
  • Sonnet 4.5
Hinglish · cost vs accuracyCheaper to the left. Training buys the top of the board.
60708090$0.01$0.10$1Stock Kev · 59.7Kev + LoRA · 88.1Jev 1.13 · 79.2Sonnet 4.5 · 81.2
Same answer, several messagesThe fine-tune stays with Sonnet or ahead on the whole test, not on one prompt.

Card has not arrived

I still have not received my new card, I ordered over a week ago.
Mera naya card abhi tak nahi aaya, maine ek week se zyada pehle order kiya tha.

ATM cash never came out

My card was denied at an ATM earlier today but the transaction is pending.
Aaj pehle ATM pe mera card deny ho gaya tha lekin transaction pending hai.

Card lost or stolen

I can't find my card and think it may have been stolen.
Mera card nahi mil raha aur mujhe lagta hai shayad chori ho gaya hai.

Refund never arrived

I requested a refund, and never received it. What can I do?
Maine refund maanga tha aur kabhi mila nahi. Main kya karun?
  • 92.7% fine-tune, English
  • 81.4% Sonnet, English
  • 88.1% fine-tune, Hinglish
  • 81.2% Sonnet, Hinglish

The job

Which queue does this customer message go to?

A bank sees this all day. The message has to be filed under the right reason: card not delivered, a refund that never arrived, cash that never came out of the ATM. There are 77 reasons in this test.

01

Sort the message

App, WhatsApp, or a call note. Send it to the right queue.

02

English is the easy case

A large general model can already do the English version.

03

Customers don't type that way

In India they mix Hindi and English. The large model is also slow and expensive.

The dataset

A public set of bank customer messages.

Banking77 is not one bank's private inbox. It is a public dataset published by PolyAI: customer messages about everyday banking, each tagged with one of 77 reasons. Source on Hugging Face.

  • 77 reasons
  • 10,003 used to train
  • 3,080 held back to grade

What we translated

Same complaint. Typed the way a customer would.

The public messages are in English. We rewrote each one as Hinglish: Hindi words written in English letters, mixed with English. Not formal Hindi, and not a word-for-word translation. The reason tag stayed the same.

Original English

My card was denied at an ATM earlier today but the transaction is pending.

Rewritten Hinglish

Aaj pehle ATM pe mera card deny ho gaya tha lekin transaction pending hai.

What it means: the ATM refused the card, the cash never came out, and the app still shows the withdrawal as pending. The model only has to pick that reason. It does not write a reply.

Leave these words as typed

  • UPI
  • ATM
  • PAN
  • FD
  • NEFT
  • OTP

Throw the rewrite out if

  • It is still plain English
  • It is only in Hindi script
  • The request changed

Messages used to grade

GrokRewrote the 3,080 held-out messages

Messages used to train

GPT-4oRewrote the 10,003 training messages

How we tested

Train a small model. Compare the bill.

  1. 01

    Start in English

    Public customer messages, each already tagged with the reason.

  2. 02

    Rewrite how people type

    Same request, same answer. Hindi in English letters, mixed with English.

  3. 03

    LoRA on a GPU

    Two passes on one NVIDIA L4. A small attachment, not a new model from scratch.

  4. 04

    Grade everyone

    Same unseen messages. Right reason, time to answer, and the bill.

LoRA

A small attachment

We kept the existing small model and trained a thin add-on. The whole model was not retrained.

One GPU

An NVIDIA L4

Two passes over the English messages and their Hinglish rewrites.

Same size

The bill does not move

Sorting stays about $15 per million messages. Training bought accuracy, not a bigger model.

Large general model

Sonnet

Given a written instruction. Accurate, and the expensive option.

Smaller product, as-is

Jev

No extra training. Cheaper than Sonnet, weaker on these messages.

Untrained

Small model, before

Fine in English. Drops to 59.7% when the message is Hinglish.

Trained on both

Small model, after

88.1% on Hinglish. Same running cost as before training.

Other frontier models

We did not run Fable. GPT-5.6 is a smaller published test.

Small model, trained

English 92.7%Hinglish 88.1%
Our test, 3,080 messages

Sonnet

English 81.4%Hinglish 81.2%
Our test, 3,080 messages

GPT-5.6

English 87.5%Hinglish Not tested
A published English test of 208 messages, not this one

Claude Fable

English Not testedHinglish Not tested
We have no score for it here

To get their score

About $20 for Fable. About $5 for GPT-5.6.

Sonnet cost about $6 because the list of 77 reasons was cached and only the short label was billed in full. Fable is $10 / $50 per million tokens against Sonnet 4.5 at $3 / $15, so the same run is about 3×, not a fresh-prompt bill.

Claude Fable

$20Same cached prompt as Sonnet. About 3× the $6 Sonnet run.

GPT-5.6

$5Terra, same setup. Cheaper per token than Sonnet.

This assumes thinking is turned off. If Fable or GPT-5.6 thinks before answering, the reply is much longer and the bill is several times higher.

A million messages

What the contact center would pay.

These two we already measured. Scale that bill to one million messages.

Sonnet

$973Measured. About 2 seconds a message. ₹81,200.

Trained small model

$15Measured. 64 milliseconds a message. ₹1,270.

Measured on a public dataset. A working comparison, not a price quote.