AI everywhere

Get AI on every job in the CRM without the runaway invoice: each feature runs on its own model, spend is metered per feature, and you set the cap.

AI everywhere

Not a module. A property of the product.

Most CRMs sell you AI as an add-on: one assistant, one model, one line on the invoice. MADDOX has more than ninety AI features, and each one is pinned to its own model in configuration — a cheap model for extracting a figure from a document, a strong one for grading a call. Every call any of them makes is written to the same usage ledger and reads back per feature. The spend cap is narrower than the ledger: it counts the assistant’s own features and brand research rather than all of them, and Cost control names the set. Changing which model does a job is a configuration change, not a release. For anyone watching margin on every seat, that is the difference between AI as a line item and AI as an unpredictable bill.

In the product

Where the money went, by feature.

The MADDOX AI cost-per-deal report: total AI spend of $12.73 across 454 calls, tiles for deals with AI spend, average per deal and the most expensive deal, a daily-spend bar chart, a spend-by-outcome table, and a ranked breakdown of what the money went on by feature.
model_pinned_features: 90+ ledger: ai_usage_logs capped_set: assistant + brand research Sample data — illustrative product UI, not a performance claim.

The controls

What an administrator actually holds

Two things decide whether AI in a business tool is safe to leave switched on: what it costs and whether it is telling the truth. Both are surfaces here, not promises.

Voice, and limits

How it sounds, and what it will not do

The last two pages are the ones a careful buyer reads first. One is about tone. The other is about the thing every vendor overstates.

Across the product

Where those ninety-plus features actually live

They are not gathered on an AI screen. They sit inside the work — in a prep document, on a call, against a deal — which is the whole argument for pinning them separately.

The machinery underneath

What all ninety-plus features share

Pinning a model per feature is the visible half. The half that decides whether any of it is safe to leave switched on is a small set of shared components every one of those features runs through, and each exists because of a specific way an AI feature fails quietly.

Nothing is ever silently truncated to fit

Every prompt component with a size budget goes through one fitter, and it has three outcomes in strictest-wins order. If the text already fits, it is passed through exactly unchanged — no marker, no extra call, byte for byte what it was. If it does not fit, the opening and closing passages are kept verbatim and only the middle is condensed by a cheap model, wrapped in a disclosure line that says which component was condensed, how much, and that the summary is explicitly not a source to quote from. If that condensing call itself fails, the text is cut and the cut is disclosed by label and by omitted character count.

What the fitter never does is drop text without saying so. That is the whole design: a model handed a quietly shortened document does not know it was shortened, and neither does the person reading what it wrote. Each outcome is also recorded as a measurement against the call it belonged to, so budget pressure is queryable before it starts costing anyone an answer.

One place decides whether a failure may be retried

A permanent failure is defined narrowly: no amount of elapsed time and no number of attempts changes the outcome, and a human has to change something. Four conditions qualify. A missing or refused provider credential. A valid credential that is out of budget — a spend wall does not clear on a backoff, and every re-attempt is refused for the same reason while still costing a request. An embedding vector that came back the wrong width, which means two settings disagree and no retry reconciles them. And a response the model truncated, which re-runs on the same input and truncates again.

Everything else is treated as transient on purpose. Connection resets, timeouts, rate limits and provider errors all keep their retry, and the default for a condition nobody anticipated is to retry rather than to discard the work. Choosing the other default would trade one defect for a quieter one. The classifier also walks the chain of wrapped exceptions, so a permanent fault buried inside a domain-specific error is still recognized as permanent.

Prompt caching is applied where it is real and skipped where it is not

Long conversations re-send a large stable prefix on every round trip, so the cacheable part is marked rather than re-billed. Two markers, deliberately: one on the block that never changes, and one placed after the last replayed turn so the cached prefix grows with the conversation instead of staying a fixed saving on a long thread. Marking the system block also sweeps in the tool definitions ahead of it, which is the largest stable thing the assistant sends. It is applied only where the model actually supports it — sending the marker elsewhere changes the shape of the request payload, and a silently rejected request on the assistant’s transport is a far worse outcome than an uncached prompt.

The ledger is wide and the cap is narrow, and it is worth knowing which is which

Every call any feature makes is written to the same usage ledger with its model, its token counts and its cost, and read back per feature, per user and per period. The spend cap is a separate and narrower thing: it counts only features whose key begins with one of a configured list of prefixes, and by default that list holds two — the assistant’s, and brand research, which is administrator-initiated spend that should not be able to run unbounded. So the ledger covers every feature and the ceiling covers those two. Widening the counted set is a configuration change rather than a code one, but it is a decision somebody has to make deliberately — and until they do, the honest statement is the one on this page. Cost control names the counted set and explains what the counter does when the store behind it goes down.

Questions

The things people actually ask.

Is our data used to train the models?

No. There is no training or fine-tuning anywhere in MADDOX — your records are never learned from. Each feature calls a commercial model API, and which provider serves which feature is configuration rather than a secret. What a provider does under its own API terms is set by that provider’s agreement, and we will tell you in writing which provider serves which feature on your workspace.

Which model does MADDOX use?

There is no single answer, and that is the point. More than ninety features each name their own model in configuration, so extracting a field from a document and grading a discovery call do not have to be billed at the same rate. Swapping a model is a configuration change. No feature hard-codes one.

What does the AI cost?

That depends on which features your people use and how much, which is exactly why every call is written to a usage ledger you can read per feature, per user and per period. What an administrator controls is the cap. See Cost control for how the cap behaves when the counter behind it goes down.

Does the AI do things by itself?

No. Every change to your data goes through a person: the assistant proposes and you approve, agents are configured rather than dispatched, and no feature in the product sends an email, creates a record or reassigns an owner without someone doing it. The Agents page says exactly where that line sits.

Can I turn AI off?

Yes. AI is a module, and modules can be switched off per tenant. When one is off the assistant stops offering the capabilities behind it rather than offering them and failing — the same predicate gates the routes and the assistant’s own capability list, so the two cannot drift apart.

How do you know the answers are grounded in my data?

A sample of real answers is checked every day by deterministic tests — does every cited record exist, are there specifics with no source behind them — and anything suspect goes into a queue for a person to rule on. Those checks make no model calls, so running them costs nothing.

Read a month of AI spend, feature by feature.

Most vendors show you one number. Open the ledger and see which feature spent what, on which model — then decide whether you would change any of it.