kaaryabench.

KaaryaBench

कार्य — the work that has to get done

Can a model finish the job when the request arrives in Hindi, Tamil, or Odia?

01

What is KaaryaBench?

A benchmark for task completion, not language quality.

KaaryaBench measures whether a language model completes a real Indian workflow when the request arrives in an Indian language. It does not measure translation quality. It measures whether the work gets done.

Why it does not exist yet

A team builds an agent that helps farmers apply for a KCC loan. The agent works by voice, in Hindi. The team tests it in English, with clean input. The tests pass.

A farmer gives his income as दस-बारह हज़ार रुपये. The form holds four income brackets. The agent selects one and does not ask which.

The same farmer gives his land as दो बीघा. The agent records two acres. Two bigha is about half an acre in his district. The record now overstates his holding by a factor of four.

The agent returned correct Hindi. It also returned a wrong record. These are two different measurements. Only the second one decides whether the application survives.

02

What should exist

One benchmark for the workflows India runs on.

  1. 01

    Real workflows

    KCC loan onboarding, PM-Kisan eligibility, and UPI transfers. Tasks drawn from Indian rails. There is no American unit called a bigha.

  2. 02

    Native evaluation

    Each task is written and graded in its own script. Nothing passes through English on the way in or the way out.

  3. 03

    Scored on the final state

    The benchmark compares the finished record against the expected schema, field by field. Fluent output that files the wrong number still fails.

  4. 04

    Two scores per model

    One score for the model alone. One score for the same model inside a production harness. The gap shows how much engineering the model needs before it ships.

  5. 05

    Ambiguity is graded

    Where the input is genuinely ambiguous, a clarifying question scores higher than a confident guess.

03

What we are working on now

Does translation already solve this?

One objection comes up every time. Translation is solved. Convert the Hindi to English, run a strong English agent, convert the answer back. We do not know yet whether that holds. We are testing it before we build the benchmark.

  1. A

    Native

    The model reads the Hindi and does the work.

    The score the leaderboard will report

  2. B

    Bridge

    A model translates the Hindi to English. An English agent then does the work.

    The objection, tested directly

  3. C

    Control

    The same tasks written in clean English.

    The ceiling. Task difficulty without language difficulty

One task pack, Hindi only. The bridge never appears on a leaderboard. It runs once. If the bridge matches the control, the objection holds and this benchmark is not needed. If the bridge loses the income bracket or the land unit on the way through, the failure is real and it is measurable.

Before it goes public

Review it before anyone trusts it.