Case Study Enterprise operations · Logistics 2026

Just ask.
Get a number you can trust.

A conversational analytics platform for last-mile operations: plain-language questions become audited, governed queries, live data, and root-cause diagnosis — for operators and managers, not just analysts.

By Javier Aragón Navarro
Built inside a multi-vertical tech group · e-commerce · logistics · fintech
80,000+employees, group-wide
3verticals: commerce · shipping · fintech
5countries in scope
Built and adopted inside a top-tier Latin American technology group spanning e-commerce, logistics, and fintech, where any production system has to earn adoption against thousands of in-house specialists.

The problem: the data existed, but only specialists could reach it

Every operational metric lived in the data warehouse and in formal dashboards. But writing a correct query meant knowing the data model, the right filters, and the canonical denominators — get one wrong and the number is misleading. So a hub manager who wanted to know whether productivity dropped this week, and why, opened a ticket and waited. The bottleneck was never the data. It was access.

Dashboards for experts

The reports needed model knowledge, correct filters, and specific calculation logic to avoid showing the wrong figure.

A warehouse out of reach

Tables run into hundreds of gigabytes in hourly partitions. A query written without the canonical rules quietly returns bad data.

No root cause

Knowing productivity fell is only the start. Explaining why took hours of manual analysis — so it rarely happened.

How it works: a question becomes an audited answer in seconds

The user asks an operational question. The platform resolves the scope, generates canonical SQL verified against a governed rules codex, runs it against the warehouse, and answers in chat with a clean table, week-over-week context, and only the glossary terms that appear.

ScopeStep 1

Resolve who's asking and about what

If the question doesn't specify a facility, operation type, or period, it asks first. It maps the user to their area of responsibility so the answer is already in their scope.

Audited SQLStep 2

Generate and audit the query against the codex

The query is built with inline references to a governed business-rules codex, then audited against a set of rules: correct calculation mode, canonical denominators, valid time filters. If a rule fails, the query never runs.

ExecuteStep 3

Dry-run for cost, then run on the right table

A dry run estimates cost first. If it passes, it executes against the correct table for the level of detail requested — facility KPIs, individual drill-down, or adoption metrics.

AnswerStep 4

Table + week-over-week + outliers + glossary

A clean table in chat. Weekly questions always include the prior week as context and flag whether a change is structural or a one-off. A short glossary covers only the terms used.

Grounded, not guessing

Every answer is anchored to one source of truth

The governed codex — canonical metrics and definitions kept as versioned documents, separate from the warehouse — defines what each KPI means and how it's calculated. The AI reads it before generating any query, so the platform's numbers match the official reports and dashboards. An AI without an anchored source of truth is a risk, not an advantage.

From "what happened" to "why"

Resolving the number is half the job. When productivity drops on a dominant sub-process, a causal model takes over and explains the drivers in plain operational language.

Causal diagnosis

A causal attribution model, auto-activated

A causal attribution model (XGBoost + OLS with SHAP) is trained on real operational history. It doesn't predict — it attributes. When there's a meaningful week-over-week drop and the facility has a reliable model, it surfaces the driver automatically: a heavier mix of large items in sorting, more staff assigned than volume justified, idle time. Outputs are always translated to operational language, never raw coefficients. If the model isn't confident for that facility, the platform shows only the observed numbers.

Representative conversation (illustrative figures)
How was net productivity at our main hub last week? Day by day.
Net productivity (units/hour) — this week vs last week
Mon 42.1 (+3.2%) · Tue 43.7 (+6.1%) · Wed 37.4 (−10.7%) · Thu 41.0 (+1.2%) · Fri 44.2 (+11.1%). Week avg 41.7 vs 40.8 (+2.2%). Wednesday breaks the pattern, but last week's Wednesday was strong too, so it's not structural — a one-off dip worth a look. Want me to break it down by sub-process or run the root-cause analysis?
Net productivity = units processed per productive hour · WoW = week over week
Yes, why did Wednesday drop?
Sorting held 68% of the hours that day with a −14.2% drop. The causal model points to a higher share of large items — the mix had ~12 points more large pieces than the week's average, which lowers throughput per hour in sorting.
Numbers above are fabricated for illustration and contain no client data.

The impact

~25 sec
to get a daily KPI with week-over-week context
vs 1–3 days through a ticket
~100
facilities with automatic causal diagnosis
across 5 countries
0
intermediaries to query operational metrics
operator, lead or manager asks directly
Built with  LLM orchestration (Claude) · BigQuery · XGBoost + OLS + SHAP · governed rules codex · MCP connectors

Want trusted self-serve analytics your team will actually use?

This is the kind of system I design and ship. Let's talk about your operation.

Get in touch