Tacnode™
Back to Blog
Architecture

HTAP May Be “Back,” But Convergence ≠ Coherence.

A decade after HTAP stalled, OLTP and OLAP are converging again — open table formats, mature CDC, engines moving to the middle. The convergence is real. But it answers where your data lives, not whether an automated decision read one coherent version of it at the instant it fired. Those are different problems.

Alex Kimball
Alex Kimball
Product Marketing
10 min read
Scattered horizontal strata sweep in from the left and collapse into a single bright band — but inside that band the lines never resolve into one line, weaving and interfering with each other instead: converged into one place, still out of register with themselves

TL;DR: HTAP — hybrid transactional/analytical processing — was declared dead somewhere around 2020. It’s back, under new names and with better foundations: open table formats, separated storage and compute, CDC that actually works, analytical engines adding transactional paths and transactional engines adding columnar ones. This round is more credible than the last one. It also answers the same question the last one did: where should the data live? That’s a placement question. The question an automated decision asks is a timing question — at the instant it fired, did it read one coherent version of everything it depended on, including the parts that had to be computed? Converging OLTP and OLAP removes copies. It does not compute your derived context, keep it fresh, or hold it coherent while thousands of concurrent decisions read and write it. That’s a separate layer, and skipping past it is how HTAP got written off the first time.

Two panels. Left, the placement question: OLTP and OLAP feeding into one slab marked ONE PLACE — convergence as a storage property. Right, the moment question: events arriving on a timeline, a stored record current at the instant the decision fires, and a derived signal still catching up behind it — coherence as a decision property.

Why HTAP Is Suddenly Everywhere Again

Convergence and coherence are different properties. Convergence is about where data lives — one system instead of five, one copy instead of three. Coherence is about what a single decision saw at the moment it executed: whether the raw record and the derived signal it read came from the same version of reality. You can fully converge your stack and still hand a decision two different moments in one query result. Convergence is a storage property. Coherence is a decision property.

On Tuesday, roughly 400 practitioners spent a day at the Rows & Columns Summit in San Francisco on a single question: what happens as transactional and analytical workloads stop being separate concerns. Andy Pavlo opened it. Hannes Mühleisen gave a talk called “Nobody Knows What OLTP Is: DuckDB Moves to the Middle.” Snowflake’s Russell Spitzer argued for Apache Iceberg as the interoperability layer between the two sides. Fatma Ozcan closed it.

That’s a serious room, and the premise is correct. The wall between OLTP and OLAP was always an artifact of hardware and storage economics, not a law of nature, and the artifact is dissolving. Databricks shipped LTAP on Lakebase in June. Analytical engines are growing transactional paths; transactional engines are growing columnar ones. Open table formats made a shared substrate plausible in a way proprietary storage never did.

The reason this deserves a careful look rather than a round of applause is that we have run this play before, and it stalled — not because the engineering was wrong, but because of what the industry decided to measure.

What HTAP Actually Promised, and Why It Stalled

HTAP was Gartner’s term, coined in 2014 for systems that run transactional and analytical workloads on one engine without ETL between them. For a few years it was the most exciting idea in data infrastructure. Then it quietly stopped being one. Search for it today and the top results include a Medium post titled “HTAP: Still the Dream, a Decade Later” and an r/dataengineering thread titled “HTAP is dead.”

OLTP vs OLAP vs HTAP, in one table.

Read the middle column down. Three categories, three answers, every one of them starting with the same word. Every one is about a workload and where it runs; not one says anything about a moment, or about what a single decision could see at the instant it had to act. That omission isn’t a flaw in HTAP — it’s the category boundary. It’s also why a converged engine can deliver everything it promises and still leave the decision problem exactly where it found it.

The usual explanation is a workload conflict: row storage serves point lookups, columnar storage serves scans, and one engine trying to do both does neither as well as a specialist. That’s true and it’s not the interesting part, because this round has real answers to it — hybrid storage layouts, separated compute, vectorized execution over open formats.

The funnier explanation is that the industry didn’t solve HTAP so much as stop saying the word. The same unfinished idea has since been sold to you as translytical, as the lakehouse, as unified analytics, as LTAP, and — this month — as engines “moving to the middle.” Twelve years, five names, one problem, and the naming is comfortably ahead of the solving.

The more useful explanation is that HTAP was sold as a consolidation story and therefore got evaluated as one. The pitch was fewer systems, no ETL, one copy. So buyers compared it on the axes a consolidation story implies: does it run my transactional workload acceptably, does it run my analytical workload acceptably, is the total cost lower. Plenty of HTAP systems could answer yes, yes, and maybe — and it still didn’t matter much, because none of those answers changed whether a decision made at 2:14:07 PM during a traffic spike was right.

Consolidation is an operations win. It shows up in headcount, in pipeline count, in the number of on-call rotations. It doesn’t show up in the outcome of any individual automated decision, which is where the money actually is. HTAP never lost an argument. It just spent a decade winning one nobody was betting on.

Round two is repeating the setup, in better clothes. Open formats, interoperability layers, engines moving to the middle — read the talk titles and the announcements and almost all of it is about where data sits and how many copies of it exist. Genuinely better answers. To the question from last time.

The question it answersBuilt for
OLTPwhere transactional work runsone record at a time, row-oriented
OLAPwhere analytical work runsmany records at once, columnar
HTAPwhere both of them runboth on one engine, hybrid storage

Convergence Answers “Where.” Decisions Ask “When.”

Here is the distinction the convergence conversation keeps sliding past.

When a fraud check, a credit authorization, a liquidation call, or an autonomous agent commits to an action, it runs inside a narrow validity window — roughly ten milliseconds to a second — with no human gate and a consequence that can’t be taken back. Outside that window the decision is already made and correction isn’t free; inside it, whatever the decision can see is the whole world. Inside that window it doesn’t read a table. It reads a bundle: an account record, a rolling velocity count over the last few minutes, an exposure total across open positions, maybe a similarity score against known-fraud patterns.

The account record is stored. Almost nothing else in that bundle is. Velocity counts, exposure totals, embeddings, risk signals — those are derived. They have to be computed from a stream of events and kept current, and the computing is where the lag lives.

This is the distinction Boyd draws in ACID for Agents as storage convergence versus computation convergence, and it’s the whole argument in one line: HTAP converges storage. Decisions run on computation.

Now put a converged system underneath that. You have removed the copy between your operational store and your analytical store, which is genuinely good. You have not:

  • computed the velocity counter — something still has to maintain it as events land
  • removed the lag between an event arriving and the derived signal reflecting it
  • guaranteed that the stored record and the derived signal in one query result describe the same moment
  • coordinated the thousands of concurrent decisions reading and writing that same state at once

Every one of those is a property of when and how derived state is maintained, not of where base data is stored. A perfectly converged stack with a derivation pipeline bolted to the side has exactly the same decision-time problem as a five-system stack with the same pipeline. The copies are gone. The moment mismatch isn’t.

Which is what makes “one copy of the data” such a convenient claim. One copy of the base data — absolutely, and it’s worth having. The derived context the decision actually leans on is a second thing, maintained somewhere else, arriving on its own schedule, and it is quietly not what anybody means when they say “one copy.” The sentence is true. It’s just answering a smaller question than the one you asked.

Four things one fraud check reads: an account record marked STORED, and three marked DERIVED — a five-minute velocity count computed from a stream of events, an exposure total aggregated across open positions, and a similarity score against known-fraud patterns. Converging OLTP and OLAP settles where the first row lives and computes none of the other three.

Three Ways a Converged System Still Hands a Decision the Wrong Context

Pipeline lag on derived state. A Flink job computes the velocity counter and writes it to Redis, while the aggregate lands in ClickHouse and the embedding goes to a vector store. Each of those advances independently of the transaction that triggered it, and of each other. Converging OLTP and OLAP doesn’t shorten that path, because the path isn’t between OLTP and OLAP — it’s between an event and a computed value. The decision reads a counter that reflects a moment already past.

Inconsistent reads across a single bundle. The decision pulls the stored record and the derived signal in one logical read. If they’re maintained by different mechanisms on different clocks, they describe different moments, and the combination never existed as a real state of the world. The decision reasons perfectly about a situation that was never true.

Incoherence under concurrency. This is the one that survives every other fix. Ten decisions touching the same account in the same second each need to see what the other nine are doing. State changes fast enough that any pre-computed view goes stale before it’s served, and concurrent load invalidates it faster than maintenance keeps up. Two decisions read two different versions, both approve, and the cap is breached. Consolidating storage does nothing here — the problem is coordination, not placement. It’s also the reason this shows up as a load problem rather than a correctness problem in most postmortems: under concurrency the same architecture that was fine at low volume stops being fine.

None of these are exotic, and they have names and fixes. They show up in ordinary language on ordinary calls: we overshoot limits under burst. We serialize per account but it kills latency. The balance doesn’t reflect the last transaction. We capped throughput to stay safe. Those sentences are all the same sentence. They describe a decision that didn’t have what it needed when it ran, and not one of them is fixed by reducing your copy count.

What Coherence Actually Requires

If convergence is a storage property, coherence has to be built where the derivation and the concurrency live. Three things have to be true at once.

Derived context is maintained inside the same system that serves it. Not computed by an external job and shipped in. When velocity counters, aggregates, and similarity signals are maintained as incrementally-updated views against the same ingested events the raw records came from, the raw and derived halves of a decision’s bundle stop describing different moments. In the Tacnode Context Lake™, that maintenance is asynchronous and converges sub-second — not zero lag, but bounded and small enough that at normal per-account rates, state propagates before the next decision on that account arrives.

Every retrieval pattern resolves against the same state. A decision needs a point lookup, a range scan, an aggregation, and often a similarity search. If those live in four systems, they pull from four snapshots and the result is a composite of four moments. Hybrid row and columnar tables let all four resolve against one internally coherent view.

Worth saying plainly, because the shape of this argument invites the assumption: this is not a claim to be HTAP done properly. HTAP is a database category, defined by which workloads one engine can run. A context lake is not a database category and doesn’t compete for that title — it’s context infrastructure, and it’s defined by what one decision can see at the moment it executes. Hybrid storage appears in both, which is why the confusion is easy, but it’s doing a different job here: not serving a transactional workload and an analytical workload from one engine, but serving one decision’s four retrieval patterns from one coherent moment. If you’re shopping for an engine to consolidate your OLTP and OLAP workloads, a context lake is the wrong product and you should buy a converged database.

Concurrent decisions are coordinated, not just served. When the state a decision writes back is owned by the context layer itself, concurrent updates are serialized under ACID and the cap-breach case is genuinely prevented. When the context layer reads from an external system of record instead, it isn’t in that write path and can’t prevent a conflicting write — what it does is make sure the decision sees accurate current context, so it doesn’t approve something it would have denied had it seen the real state. Worth being precise about which of those two you’re being sold.

Where Convergence Genuinely Wins

It would be dishonest to take the summit’s premise apart without saying what it gets right, because it gets a lot right.

If you maintain three copies of the same table and spend real time reconciling drift, convergence attacks that directly and the win is immediate. If your analytical queries run against data that’s a day old because the pipeline runs nightly, convergence collapses that to something far better. If your governance model has to be re-implemented per system, a shared substrate under open table formats is a straightforward improvement. If you’re staffing a pipeline team whose entire job is moving rows between engines that should have shared storage, that’s a cost convergence removes.

Those are real problems and a converged architecture is the right answer to them. The claim here is narrow: those are data platform problems. They’re upstream of decision time, they’re solved by a different layer than the one that makes an automated decision correct under load, and the two layers sit in the same stack without competing.

The failure mode isn’t adopting convergence. It’s adopting it and assuming the decision-time problem came along in the box.

The Question to Ask Any Converged System

One question separates the two properties, and it’s worth asking of any vendor, any architecture diagram, and any conference talk premise:

When a decision reads a stored record and a derived signal in the same request, are both computed from the same set of events — and does that still hold when a thousand of those decisions run against the same state in the same second?

If the answer is about copy count, storage layout, or governance, you got an answer to the placement question. It may well be a good answer. It is not an answer to this one.

HTAP round one was a good idea measured on the wrong axis, and it spent a decade in “still the dream” posts as a result. Round two has much better technology and exactly the same framing. The engineering will probably hold this time. Whether anyone remembers it in 2036 depends on whether the question moves from where does the data live to what did the decision see — and right now the conference programs suggest it hasn’t.

FAQ

HTAPOLTPOLAPDecision CoherenceArchitecture
Alex Kimball

Written by Alex Kimball

Former Cockroach Labs. Tells stories about infrastructure that actually make sense.

Fresh Context Newsletter

The argument behind the posts — what breaks when AI acts on stale data, from the team building the Context Lake. About twice a month.

Subscribe on LinkedIn

Ready to see Tacnode Context Lake™ in action?

Book a demo and discover how Tacnode can power your AI-native applications.

Book a Demo