Agentic systems decide at run time. Control has to live there.
Kalmi builds the layer where policy is enforced, spend is bounded, and every action becomes evidence. For banks, NBFCs, and the India entities of global financial institutions.
The pilot works. The system still does not ship.
Some pilots fail on accuracy, on data that was never as clean as the demo, or on a case that did not survive a real cost line. This is about the other ones: systems that work in test, have a sponsor and a budget, and still do not reach production.
The pattern is consistent. The policy exists as a document no runtime reads. The logs were built for an engineer debugging a failure, not for a reviewer reconstructing a decision a year later. The controls are described in a framework and unowned in production. The inventory is a spreadsheet that becomes accurate shortly before an audit.
This is not an absence of controls. The controls matrix is real and most of it works. It was built for systems that decide once, at release, and behave the same way until the next release. An agent chooses its own actions long after that gate has closed.
Risk and compliance are not being obstructive when they decline to sign. They are being asked to accept a system they cannot inspect, constrain or evidence.
Three things have to exist before an agent runs in production.
None of them are produced by a policy review. Each one is built.
A model inventory that is current by construction
A register the systems populate rather than one a person updates before an examination: every model and agent in use, its owner, its version, where it runs, what data it touches. An inventory maintained by hand is wrong on the day it is read, and everything downstream inherits the error.
Policy enforcement in the runtime path
Each obligation expressed as a rule the execution path evaluates before the action completes. A data category that may not leave a jurisdiction. An action that requires approval above a threshold. A model not cleared for the use case. The control fires or it does not, and either outcome is recorded.
Audit logging built to an evidentiary standard
Application logs answer why something broke. Evidence answers what the system decided, on which inputs, under which policy version, and who approved the exception. The second cannot be reconstructed from the first, which is why the evidence model is designed before the first log line is written.
Four tiers. Each one is saleable on its own and each one earns the next.
Most engagements start at the first. Nothing here requires buying the tier above it.
Diagnostic
Two to three weeks, fixed fee, scoped before it begins, remote. Current state, regulatory exposure, a triaged use case portfolio, a roadmap.
One rule separates it from the assessments you have read. It does not conclude without naming who operates each control it recommends.
Remediation
Four to eight weeks, scoped after the diagnostic. Model inventory, governance policy, a controls matrix mapped to named obligations, accountability structure, documentation standards.
Every control is written so a runtime can execute it. A control expressed only as prose gets rewritten the day you decide to enforce it.
The control plane
The policy set derives from your obligations rather than a template, and the evidence model is designed before the first log is written. Deployment defaults to your environment, under your controls and identity management, which keeps residency simple and exit clean.
Gateway, routing, caching and budget control are commodity. We configure that layer rather than compete with it.
Operated workflows
Specific workflows run in production on the control plane, each sold by name with its own scope, its own controls, a defined support window, and exit terms agreed at the start.
Three shapes we are set up to build:
Periodic KYC review. Accounts approaching re-verification are monitored, documents requested and validated, exceptions routed to a reviewer. Every automated decision and every override is recorded as evidence.
Regulatory change mapping. New circulars, directions and parent group standards are mapped to the internal policies and controls they affect, with the revision drafted for a named owner to approve.
Pre-submission reporting review. A regulatory or bureau submission is checked against the applicable requirements before filing, with anomalies flagged for correction rather than discovered by the recipient.
A workflow that cannot be handed back is a dependency your risk function will object to, so the handback path is designed in at the start.
What production-ready means here.
A system is production-ready when it delivers commercial return while meeting all six of these, defensible under the frameworks that apply to your institution. Four out of six is a pilot with a launch date attached.
Explainability. A decision can be reconstructed and explained to someone who was not there, in the language the business uses.
Transparency. What the system does, on what data, under which policy version, is visible without reading the code.
Consistency. The same inputs receive the same treatment, and drift is detected rather than discovered.
Correctness. Accuracy is measured against a defined standard continuously, not demonstrated once at launch.
Bias management. Assessed on the attributes that matter for the use case, with a documented method and a retest cadence.
Commercial return. The case still holds after the cost of running the system under the five constraints above.
This is the standard we hold client systems to, and the one our own work gets measured against.
Where we sit, and what we do not sell.
First line of defense
Kalmi builds controls and operates them, which places us in the first line. We do not provide second-line validation or third-line assurance on our own work, because the party that builds a control cannot challenge it. Your model risk function validates what we build. We build so that it can be.
No commercial interest below the control plane
No vendor relationships, no resale, no referral fees, no preferred model provider, no preferred cloud. Nothing we recommend beneath the control plane earns us anything.
Opinionated at exactly one layer
Neutrality below the control plane is a commercial fact, not an architectural position. We have a view on how the control layer itself should be built, and we state it rather than charging you for the time it takes to choose.
We decline work
Most of an AI roadmap should not be built right now. Every diagnostic names the use cases that should not proceed and says why. We hold ourselves to the same filter: one product, a bounded set of named workflows, two segments.
Sudip Chakraborty
Sudip was a Vice President at Goldman Sachs, working on stress testing regulation, board-level governance and buy-side market risk: the discipline of making a model defensible to a supervisor rather than merely accurate. At Soroco he implemented ISO 27001 and SOC 2 compliance, which is control work rather than framework theory. IIT Guwahati, IIM Kozhikode.
This work requires the technical, the domain and the regulatory at the same time. Teams strong in two of the three produce frameworks that collapse in production, or systems that work and cannot be defended.
Two kinds of institution, two different reasons this stalls.
The India entity of a global financial institution
The build happens here and the model risk policy it has to satisfy was written elsewhere. A pilot clears internal review, reaches the edge of production, and stops because no one can say who operates the controls once it is live.
An NBFC or bank supervised by the RBI
The examination question is not whether a policy exists. It is whether you can produce the inventory, the control and the evidence on the day you are asked, for a system that has been running for a year.
The first conversation is a diagnosis, not a pitch.
Bring a system that stalled, or an inventory you are not confident in. We will tell you what we typically see at that stage, specifically enough that you can confirm it or correct it. If an engagement follows it starts with the diagnostic, which is paid, fixed fee and scoped before it begins. We do not run free assessments.
Replies come from Sudip, usually within a working day.