Hinglish accuracy
59.7+2888.1
Stock Kev to two LoRA epochs. Sonnet sits at 81.2.
Case study · sorting customer messages
A small model sorts bank customer messages more accurately than Sonnet, at 64× lower cost.
Hinglish accuracy
59.7+2888.1
Stock Kev to two LoRA epochs. Sonnet sits at 81.2.
Cost per 1k
$0.9764×$0.015
Trained Kev on an India L4 versus billed Sonnet 4.5.
Latency
2.0s31×64ms
p50. Choice prefill versus a prompted decode.
Card has not arrived
I still have not received my new card, I ordered over a week ago.
Mera naya card abhi tak nahi aaya, maine ek week se zyada pehle order kiya tha.
ATM cash never came out
My card was denied at an ATM earlier today but the transaction is pending.
Aaj pehle ATM pe mera card deny ho gaya tha lekin transaction pending hai.
Card lost or stolen
I can't find my card and think it may have been stolen.
Mera card nahi mil raha aur mujhe lagta hai shayad chori ho gaya hai.
Refund never arrived
I requested a refund, and never received it. What can I do?
Maine refund maanga tha aur kabhi mila nahi. Main kya karun?
The job
A bank sees this all day. The message has to be filed under the right reason: card not delivered, a refund that never arrived, cash that never came out of the ATM. There are 77 reasons in this test.
App, WhatsApp, or a call note. Send it to the right queue.
A large general model can already do the English version.
In India they mix Hindi and English. The large model is also slow and expensive.
The dataset
Banking77 is not one bank's private inbox. It is a public dataset published by PolyAI: customer messages about everyday banking, each tagged with one of 77 reasons. Source on Hugging Face.
What we translated
The public messages are in English. We rewrote each one as Hinglish: Hindi words written in English letters, mixed with English. Not formal Hindi, and not a word-for-word translation. The reason tag stayed the same.
Original English
My card was denied at an ATM earlier today but the transaction is pending.
Rewritten Hinglish
Aaj pehle ATM pe mera card deny ho gaya tha lekin transaction pending hai.
What it means: the ATM refused the card, the cash never came out, and the app still shows the withdrawal as pending. The model only has to pick that reason. It does not write a reply.
Leave these words as typed
Throw the rewrite out if
Messages used to grade
GrokRewrote the 3,080 held-out messagesMessages used to train
GPT-4oRewrote the 10,003 training messagesHow we tested
Public customer messages, each already tagged with the reason.
Same request, same answer. Hindi in English letters, mixed with English.
Two passes on one NVIDIA L4. A small attachment, not a new model from scratch.
Same unseen messages. Right reason, time to answer, and the bill.
LoRA
One GPU
Same size
Large general model
Smaller product, as-is
Untrained
Trained on both
Other frontier models
Small model, trained
Sonnet
GPT-5.6
Claude Fable
To get their score
Sonnet cost about $6 because the list of 77 reasons was cached and only the short label was billed in full. Fable is $10 / $50 per million tokens against Sonnet 4.5 at $3 / $15, so the same run is about 3×, not a fresh-prompt bill.
Claude Fable
$20Same cached prompt as Sonnet. About 3× the $6 Sonnet run.GPT-5.6
$5Terra, same setup. Cheaper per token than Sonnet.This assumes thinking is turned off. If Fable or GPT-5.6 thinks before answering, the reply is much longer and the bill is several times higher.
A million messages
These two we already measured. Scale that bill to one million messages.
Sonnet
$973Measured. About 2 seconds a message. ₹81,200.Trained small model
$15Measured. 64 milliseconds a message. ₹1,270.Measured on a public dataset. A working comparison, not a price quote.