ClickHouse vs Doris vs StarRocks: OLAP ADR 2026
Last Updated: October 4, 2026
Most teams choose a real-time analytics engine from a benchmark chart and regret it a quarter later, when the first upsert workload, the third dimension table, or the thousandth concurrent dashboard user arrives. The decision of ClickHouse vs Doris vs StarRocks is no longer a single-table speed contest. It is a decision about join strategy, mutation semantics, tenancy, and how many moving parts your team can responsibly own for years.
It matters now because the three engines have converged on checklists while still diverging on failure modes. ClickHouse 26.8 shipped a cost-based distributed planner in experimental form and a new inequality-join algorithm. Apache Doris 4.1 and StarRocks 4.1 both pushed deep into Iceberg v3, shared-data storage, and search. The feature pages now look the same, which is exactly when architecture decisions get made badly.
This post is a working architecture decision record (ADR). You leave with the decision drivers, a mechanism-level explanation of where each engine is strong and where it breaks, a weighted matrix you can re-weight, a decision tree, and a migration-safe recommendation. Every number is either sourced or labelled illustrative.
What this covers: the architecture of each engine, joins and planners, mutable data and upserts, concurrency and tenancy, lakehouse federation, operations and licensing, a decision matrix and tree, and the failure modes vendor pages omit.
What Changed for October 2026
This is a ground-up rewrite of the June 2026 version of this article. If you read that one, these are the corrections and updates that matter.
- Current versions. ClickHouse 26.8 was released on September 10, 2026 and is a long-term support (LTS) release. Apache Doris is on the 4.1 line (4.1.0 on April 21, 2026 per the release page; 4.1.4.1 on September 29, 2026), with 4.0.x still maintained (4.0.8 on August 14, 2026). StarRocks has both a 4.0 line (4.0.16 on September 30, 2026) and a 4.1 line (4.1.6 on September 30, 2026).
- ClickHouse’s join story changed. The earlier post called multi-table joins ClickHouse’s soft flank. That is still directionally true, but 26.8 added column statistics collected on INSERT for small tables by default, an IEJoin algorithm for inequality joins, a parallel full-sorting-merge join, and an experimental Cascades optimizer that chooses between broadcast, replicated and shuffle joins.
- Doris is no longer just an OLAP engine. Doris 4.0 and 4.1 added vector search, a BM25-scored
search()function, AI functions, and full Iceberg v2 and v3 read and write. Whether you want that in the same engine as your telemetry store is now an architectural question in its own right. - StarRocks moved on tablets and schema evolution. StarRocks 4.1 added range-based tablet auto-splitting for multi-tenant tables, larger tablets in shared-data mode, and Fast Schema Evolution v2.
- Scores were softened. The old weighted matrix produced a winner with false precision. This version scores on a coarser three-level scale and tells you to re-weight it.
Context and Background
Real-time OLAP sits between the system that records events and the dashboard, API, or agent that queries them. In an IoT or digital twin deployment, that means machine telemetry, device state changes, and alarm streams landing in a store that must answer “what is the vibration trend for these 400 assets over the last six hours” in under a second, for many users at once. If you are mapping that problem end to end, our industrial IoT time-series platform architecture guide covers the ingestion side that feeds these engines.
The incumbents are three columnar, vectorized engines. ClickHouse began as a single-binary scan machine at Yandex and grew a distributed mode and a managed cloud. Apache Doris is an Apache Software Foundation top-level project built around a frontend and backend split. StarRocks started from a Doris lineage and rebuilt the query engine around a cost-based optimizer and a pipeline execution model; it is published under the Apache License 2.0 and hosted under the Linux Foundation. ClickHouse core is also Apache 2.0. Licensing is therefore a tie-breaker, not a driver, but managed-cloud terms differ and are worth reading.
What the three share is more instructive than what separates them. All store data by column, compress per column with type-aware codecs, skip irrelevant data using sparse indexes and zone maps, and execute filters and aggregates in vectorized batches across cores. The differences are in the control plane and the planner. ClickHouse historically gave you a fast executor and expected you to be the optimizer: you ordered the tables, you picked the algorithm, you chose the sort key. Doris and StarRocks gave you a cost-based optimizer that rewrites and reorders for you, and paid for it with more node roles.
That gap is the thread running through this ADR, and it is narrowing from one side. ClickHouse is acquiring planner intelligence while the MPP engines are acquiring cheap, elastic storage. For readers who want the incremental view of ClickHouse itself, our look at ClickHouse 26.8 LTS versus 26.3 covers what changed between the last two LTS lines. For the canonical behavior of MergeTree, use the ClickHouse documentation.
Decision drivers
An ADR is only as honest as its drivers. Five forces shaped this comparison, in priority order.
- Query shape. The ratio of single-table scans to multi-table joins in your real workload. Nothing splits the field more cleanly.
- Write pattern. Append-only event streams versus mutable rows that need correct, low-latency upserts and deletes.
- Concurrency and tenancy. A handful of analyst queries versus thousands of small user-facing requests per second, possibly from many tenants sharing tables.
- Lakehouse posture. Whether data already lives in Iceberg, Paimon, or Hive and must be queried in place instead of copied.
- Operational budget. How many node roles, tuning axes, and upgrade procedures your team can own.
Cost and licensing are tie-breakers. A thesis runs through the rest of this document: the question is not “which engine is fastest” but “which engine’s worst day is one I can live with”. Each engine has a characteristic worst day, and the sections below name it.
Scope and what was rejected
This ADR is scoped to three distributed column stores. A cloud warehouse such as Snowflake or BigQuery was set aside because the requirement is sub-second latency at high concurrency, where per-query overhead and pricing do not fit. Pinot and Druid are strong for append-only streams but weaker at ad hoc joins; they are covered in our ClickHouse vs Druid vs Pinot ADR and the companion Pinot vs Druid ADR. Embedded engines target single-node analytics, which we cover in DuckDB vs ClickHouse. Pure time-series engines are compared in our Postgres 18 vs TimescaleDB vs ClickHouse post.
Reference Architecture: How Each Engine Moves Data
Short answer. ClickHouse is a symmetric engine: every node stores parts and plans queries, and you scale by sharding. Doris and StarRocks are asymmetric MPP engines: frontends plan and hold metadata, backends (or compute nodes) store and execute, and a cost-based optimizer decides how data moves between them. The first minimizes roles, the second minimizes manual planning.

Figure 1: Control-plane and data-plane layout of the three engines, showing where planning and storage live.
Figure 1 compares the tiers. On the left, a ClickHouse cluster is a set of identical nodes, each owning shards of MergeTree tables, coordinated for replication by ClickHouse Keeper. A Distributed table is a routing layer that fans a query out and merges partial results on the initiator. On the right, Doris and StarRocks separate a small, Raft-style replicated frontend group, which parses SQL, holds catalog metadata and runs the optimizer, from a horizontally scalable tier of backends that scan tablets and exchange data between plan fragments. In shared-data mode, the backends become cache-bearing compute that reads from object storage.
ClickHouse: MergeTree, sparse indexes, and the symmetric cluster
ClickHouse stores each insert as an immutable part, sorted by the table’s ORDER BY key. The sort key doubles as a sparse primary index: one index entry per granule (8,192 rows by default) rather than per row. At trillions of rows the index still fits in memory, and a predicate on the leading key columns lets the engine discard whole granules without reading them. Background merges compact small parts into larger ones, and this merge process is where rollup and dedup semantics live, through the ReplacingMergeTree, SummingMergeTree and AggregatingMergeTree variants.
Column codecs matter more than most comparisons admit. Delta and DoubleDelta encodings suit monotonic timestamps, Gorilla suits slowly varying floating-point sensor values, and ZSTD sits on top. For IoT telemetry, where a reading is often nearly identical to the previous one, per-column codec choice is a larger cost lever than any engine selection.
The symmetric design is the source of ClickHouse’s operational appeal. A three-node cluster with replicated tables is operable by one engineer. The cost is that distribution is explicit: you pick the sharding key, you decide whether a join’s right side is small enough to broadcast, and you set the join algorithm. Our earlier claim that ClickHouse treats the planner as your job is the most important sentence in this ADR, and 26.8 is the first release where the project credibly starts to take it back.
Doris and StarRocks: frontend, backend, and the table models
Both engines parse SQL on a frontend, plan with a cost-based optimizer, split a plan into fragments, and execute fragments on backends that exchange data over the network. Both expose explicit table models. A duplicate model keeps every raw row and ingests fastest. An aggregate model pre-aggregates at load time. A unique or primary-key model supports upserts and deletes by key. The primary-key model is the one that changed the field: it maintains a persistent key index and applies delete markers at read time, so reads stay close to append-only speed instead of requiring a merge-on-read penalty.
Doris and StarRocks diverge in implementation details more than in shape. StarRocks leans harder into a pipeline execution engine and a query cache aimed at high concurrency. Doris leans into breadth: inverted indexes, full-text search with BM25 relevance, vector indexes, and AI functions in the same SQL surface. Both support tiered or shared-data storage on S3-compatible object stores, with a local cache.
Shared-nothing versus shared-data
Classic deployment places data on local disks attached to backends. That gives the lowest latency and couples storage to compute: to add capacity you add machines, and rebalancing moves data. Shared-data mode puts primary data on object storage and uses local disks as cache. You can then scale compute for a reporting window and scale it back without moving terabytes, and you pay object-storage rates for cold data. The price is cold-read latency and cache-management complexity, which is why both vendors keep investing in cache observability and warmup. Doris 4.0.7 added table-level filtering for event-driven cache warmup, and StarRocks 4.1 added end-to-end cache metrics from cluster level down to individual queries. ClickHouse offers object-storage-backed tables as well, and its managed cloud separates storage and compute, but for self-managed deployments the shared-data experience in the MPP engines is, in our assessment, further along.
Joins and Planners: The Real Fault Line
Short answer. Single-table scan speed is a three-way tie for practical purposes. Joins are the fault line: StarRocks and Doris plan distributed joins by cost, choosing broadcast, shuffle, or colocated execution automatically; ClickHouse historically left that choice to you and has just begun adding a cost-based distributed planner, still experimental in 26.8.

Figure 2: How a cost-based planner chooses a distributed join strategy, and where the memory cliff sits.
Figure 2 traces the decision. If the build side is small, the planner broadcasts it to every node and joins locally. If both sides are large, it shuffles both by the join key so matching rows meet on one node. If both tables were colocated by that key at load time, the join runs with no network exchange at all. Runtime filters then push a compact summary of the build side into the probe-side scan to discard rows early. At the bottom of the diagram sits the cliff: if the build side’s hash table exceeds memory, the engine must spill or fail.
Why this matters for star schemas
A digital twin query typically joins a wide fact table of telemetry to dimension tables of assets, sites, and model metadata, and sometimes to a slowly changing hierarchy table. Four or five joins in one query is normal. A cost-based optimizer that reorders those joins and chooses broadcast for the small dimensions and shuffle for the large ones is worth far more than a faster scan. On this shape, StarRocks and Doris have been the safer choice for years. Colocation deserves special mention: declaring that the telemetry and asset-state tables share a distribution key turns the biggest join into a local operation. ClickHouse can approximate that with matching sharding keys and distributed_product_mode-style settings, but you are the one holding the invariant.
What ClickHouse 26.8 actually changed
The 26.8 release material reports four relevant changes, and we cite its own numbers rather than inventing ours. First, statistics collection on INSERT for small tables is enabled by default, which feeds join planning; the release presentation reports a 29 percent improvement across its benchmark set and a 4.5x improvement on TPC-H. Second, the IEJoin algorithm handles inequality joins, such as matching trades to quotes by time range; the presentation shows a query dropping from 112 seconds to 1.06 seconds. Third, a parallel full-sorting-merge join ran a 100 million by 150 million inner join in 1.5 seconds on 32 threads versus 5.9 seconds single-threaded. Fourth, an experimental Cascades optimizer estimates cardinalities and selects among broadcast, replicated, and shuffle joins for distributed queries.
Treat these as vendor-reported results on vendor-chosen queries. The IEJoin number in particular is a best case, because inequality joins were pathological before. The structural point stands independent of any figure: the planner gap is closing, but “experimental” is the operative word for the piece that matters most in a cluster. A conservative team should wait for Cascades to leave experimental status before betting a multi-join workload on it.
What the MPP engines still own
Doris and StarRocks have run cost-based distributed planning for years, so the failure modes are known and the tooling exists: EXPLAIN, query profiles, skew detection. StarRocks 4.1 adds improved skew-join detection and recursive CTEs. Doris 4.1 reports aggregation pushdown through joins and grouping-set optimizations. The StarRocks 4.0 announcement claims roughly 60 percent faster performance year over year on its test sets and a 1.6x speedup on a TPC-DS 1 TB run; Doris 4.1 reports TPC-H improved 22.6 percent over 4.0. These figures are vendor-reported, run on vendor hardware, and cannot be compared with each other.
Deeper Analysis: Mutations, Concurrency, and Tenancy
Short answer. If rows change after they land, StarRocks and Doris primary-key tables give read-time-consistent upserts at near append speed. ClickHouse can do the same job with ReplacingMergeTree or lightweight updates, but you must understand merge timing. For user-facing concurrency, the MPP engines ship purpose-built caches; ClickHouse can serve high QPS, but you design the data model to make that cheap.

Figure 3: How a changed row becomes visible to a reader in each engine’s write path.
Upserts and deletes
Figure 3 follows a changed row. In a StarRocks or Doris primary-key table, a write lands in an in-memory structure and consults the persistent primary-key index to locate the previous version, then records a delete marker against it. Readers apply the marker while scanning, so they see one version. Compaction later removes the old row physically. The consequence is predictable read latency and correct results immediately, at the cost of memory for the key index and write amplification during compaction. The StarRocks 4.0 notes mention compaction optimizations, delete-vector metadata handling, and partial-update protections against schema drift, which tells you where the maintenance effort goes.
ClickHouse’s classic approach is ReplacingMergeTree: a new row with the same sort key is simply another part, and duplicates collapse when parts merge. Between merges a reader may see both versions unless you add FINAL, which dedups at query time at a cost, or write queries that aggregate by key. Lightweight updates and deletes reduce the pain for corrections and compliance deletes, but they do not turn ClickHouse into a transactional store. For append-mostly telemetry this is irrelevant. For asset-state tables that change constantly, such as “current status of each device”, it is the single most common source of surprise.
A practical rule: if more than a small fraction of your rows are updated after landing, or if readers cannot tolerate seeing a stale duplicate for seconds to minutes, use a primary-key engine. If your table is append-only and you handle the rare correction as a batch, ClickHouse’s simpler model wins.
Concurrency for user-facing analytics
User-facing analytics means hundreds or thousands of small, fast queries per second, not a few huge scans. The risk is tail latency: a handful of large queries starve the many small ones. StarRocks’s pipeline execution model schedules work as fine-grained drivers across cores rather than one thread per fragment, which keeps latency steadier as concurrency climbs, and its query cache reuses partial aggregation results across similar queries. Doris offers its own vectorized pipeline and caching; Doris 4.0 made the SQL cache default-on and reports a large improvement in SQL parsing efficiency for repeated statements.
ClickHouse serves high QPS well when the design cooperates. The pattern is to pre-aggregate with materialized views into AggregatingMergeTree targets, narrow the schema, set per-user quotas and concurrency caps, and keep heavy ad hoc work on a separate replica set. It works, and many large deployments run it, but you are doing the engineering the MPP engines partly automate.
Tenancy and tablet management
Multi-tenant telemetry, where one table holds many customers with skewed volumes, stresses data placement. StarRocks 4.1 introduced range-based tablet auto-splitting along sort-key ranges so a hot tenant’s data subdivides without schema changes or re-ingestion; the vendor reports a 1.86x throughput improvement and P99 latency falling from 36.6 to 11.5 seconds in its benchmark. Treat those as best-case vendor figures. In ClickHouse you handle skew with the sort key and partitioning scheme, which is flexible but static. Doris relies on bucketing choices made at table creation, so mis-sized buckets are a classic early mistake.
Materialized views
Doris and StarRocks can transparently rewrite a query to read a materialized view it never named, so adding a rollup later accelerates existing dashboards without changing SQL. StarRocks 4.1 extends incremental materialized views with more aggregate functions and reports 7 to 30 times faster refresh on its Iceberg benchmark. ClickHouse materialized views are insert-time triggers into a target table; they are extremely efficient but are not transparent rewrites, so the dashboard must query the rollup explicitly. ClickHouse 26.8 added atomic materialized view population, which fixes a long-standing backfill race.
Lakehouse, Search, and Ingestion
Short answer. All three can now query Iceberg in place, and all three are racing on Iceberg v3. Choose on write support and catalog integration, not on whether a connector exists. For ingestion, all three reach seconds-fresh data, but they ask for different discipline: ClickHouse wants large batches, the MPP engines want you to watch compaction.
Iceberg and external tables
A hot telemetry store rarely holds everything. Cold history belongs in object storage as Iceberg tables that Spark, Flink, and the OLAP engine all read. If you are deciding which table format to standardize on, see our Apache Iceberg v3 upgrade guide for what v3 changes.
On this axis the three engines read as follows, according to their own release material. Doris 4.1 claims full Iceberg v2 and v3 read and write. StarRocks 4.1 reports Iceberg v2 and v3 enhancements including native SQL DELETE and Variant type support, with 4.0 having added hidden partitions and a compaction API. ClickHouse 26.8 added S3 Tables write support, Snowflake Horizon catalog integration, and Puffin file support for Iceberg metadata. The honest summary is that capability parity is close, but the three are not equally mature on every catalog type and every v3 feature. Before committing, test your actual catalog (REST, Glue, Hive, or Polaris-style) with your actual write path, because “supports Iceberg” covers a wide range of behavior.
ClickHouse 26.8’s Parquet work is worth noting for lake reads: the release material reports lazy materialization that cut I/O on a tested query from 6.7 GB to 1.5 GB, about 77 percent. Again, a single query, but it shows the mechanism: read the filter columns first, then fetch only the surviving rows of the other columns.
Search and vectors in the same engine
Doris 4.0 and 4.1 pushed search aggressively: a search() function with an Elasticsearch-style query syntax and BM25 scoring, vector indexes, and AI functions that call language models from SQL. Doris 4.1 reports IVF and IVF_ON_DISK vector indexes, quantization that reduces index memory by a factor of four to eight, and single JSON documents up to 100 MB. StarRocks 4.1 lists inverted indexing in beta. ClickHouse has its own text index work, and 26.8 added Japanese, Chinese and ICU-based tokenizers.
For IoT readers, the relevant question is whether you want log search, device-manual retrieval, or agent memory colocated with telemetry. Consolidation reduces systems to run, but it couples the upgrade and failure domains of workloads that rarely need to share fate. Our view: put search in the same engine only if the same people operate both workloads and the data genuinely joins; otherwise keep the telemetry store lean.
Ingestion paths
ClickHouse favors large batched inserts straight to MergeTree parts. Each insert creates a part, and too many small parts causes merge pressure, which is why async inserts and the Kafka table engine exist: they accumulate small writes into block-sized parts. Freshness is excellent once a part lands. Doris and StarRocks provide stream load over HTTP, routine load from Kafka, and primary-key upserts that make the row visible almost immediately while compaction catches up in the background.
For a fleet of devices posting every few seconds, the practical architecture is the same for all three: put a log such as Kafka or NATS in front, let a connector batch rows, and keep the OLAP engine’s insert rate in the range of tens of batches per second, not thousands of tiny writes. If you are choosing the broker, see our comparison of NATS JetStream vs Kafka for edge IIoT telemetry.

Figure 4: Reference ingestion path for telemetry into any of the three engines, with an Iceberg cold tier.
Figure 4 shows the shape we recommend regardless of engine. Devices publish to an edge broker; a bridge forwards to a durable log; a loader batches and writes to the hot OLAP store; a separate retention job compacts older partitions into Iceberg on object storage, so the hot tier stays small and cheap. The OLAP engine then serves dashboards and APIs from the hot tier and can federate to Iceberg for long-range queries. This is also where idempotency matters: batch loaders retry, so every row needs a deterministic key or a dedup strategy that matches your engine, such as ReplacingMergeTree in ClickHouse or a primary-key table in Doris and StarRocks.
Operating Each Engine
Short answer. ClickHouse has the smallest conceptual footprint. Doris and StarRocks have more roles but more automation around placement, balancing, and tablet repair. The cost of the MPP engines is a larger vocabulary; the cost of ClickHouse is that more decisions are yours.
Footprint and failure domains
A ClickHouse cluster has one process type plus Keeper. Rolling upgrades are per-node, and monitoring is mostly system tables. The failure that surprises teams is not a crash but a merge backlog: too many parts, slow merges, rejected inserts. Doris and StarRocks have frontends and backends, so you run and monitor two roles, size tablets and buckets, and watch tablet balance. In exchange, the engine repairs under-replicated tablets and rebalances on its own, which in ClickHouse is a manual resharding exercise.
Release cadence differs too. ClickHouse ships monthly releases and designates some as LTS; 26.8 is LTS, so conservative teams can sit on it. StarRocks patch cadence is high, with 4.0.16 and 4.1.6 both released on September 30, 2026, which signals active maintenance and also a lot to track. The StarRocks 4.1 release notes carry an explicit warning that container users should avoid v4.1.0 because of an unstable load-order issue affecting backend startup. That is the kind of caveat that argues for waiting a few patch releases on any new minor line. Doris keeps 4.0.x and 4.1.x lines in parallel, which gives you a conservative and a current track.
Managed options and total cost
All three have managed offerings: ClickHouse Cloud, VeloDB for Doris, and StarRocks-based services including CelerData. Pricing models differ and change often, so we do not quote figures. The structural point is that managed ClickHouse is the most widely available option across clouds, which matters if your platform team is small. Self-hosting on Kubernetes is feasible for all three through operators, but the MPP engines’ stateful roles make storage class and cache sizing more consequential than for a symmetric ClickHouse cluster.
Decision Matrix
The matrix scores each engine on a deliberately coarse scale: Strong, Adequate, or Weak for a typical real-time analytics product on a multi-node cluster in October 2026. These are the author’s qualitative judgments, not measurements, and they will not survive a workload that differs from the one assumed.
| Decision axis | Weight (illustrative) | ClickHouse 26.8 | Apache Doris 4.1 | StarRocks 4.1 |
|---|---|---|---|---|
| Single-table scan and aggregation | High | Strong | Strong | Strong |
| Multi-table distributed joins | High | Adequate, improving | Strong | Strong |
| Mutable data and upserts | Medium | Adequate | Strong | Strong |
| High-concurrency user-facing queries | Medium | Adequate with design work | Strong | Strong |
| Ingestion of wide append-only streams | High | Strong | Strong | Strong |
| Iceberg read and write | Medium | Adequate to Strong | Strong | Strong |
| Search and vector workloads | Low | Adequate | Strong | Adequate |
| Operational simplicity | Medium | Strong | Adequate | Adequate |
| Managed-cloud availability | Medium | Strong | Adequate | Adequate |
How to use it: replace the weights with your own, count Strong as 3, Adequate as 2, Weak as 1, and sum. A single-table observability store should weight scan and ingestion heavily and will likely land on ClickHouse. A user-facing product built on star-schema joins and mutable state will land on StarRocks or Doris. The earlier numeric matrix produced a winner by summing guesses; a coarse scale is more honest about what we know.
Two cells deserve defense. “Adequate, improving” on ClickHouse joins reflects that 26.8 materially advanced the planner but the distributed cost-based piece is experimental. “Adequate” on Doris and StarRocks operations reflects node roles and tablet sizing, not instability.
Choosing between Doris and StarRocks
The two are closer to each other than either is to ClickHouse, so the tie-break is secondary. Pick Doris if you value breadth in one engine (search, vectors, AI functions), the Apache governance model, or the managed VeloDB route. Pick StarRocks if your priority is high-concurrency serving with its query cache and pipeline engine, multi-tenant tablet management, or the CelerData route. Run your own queries on both; the difference between them on your data will be smaller than the difference between either and a poorly chosen table design.
Trade-offs, Gotchas, and What Goes Wrong
Each engine has a characteristic worst day. Know it before you buy.
ClickHouse: the join and merge cliffs. The worst day is a multi-join query whose right side does not fit in memory, or an insert pattern that creates too many small parts. Both are avoidable with discipline: denormalize hot dimensions into the fact table or into dictionaries, keep inserts batched, and watch the parts count per partition. The risk is organizational: the discipline lives in people’s heads. ClickHouse dictionaries deserve special mention for IoT, because an asset-to-site lookup held as an in-memory dictionary removes a join entirely.
Doris and StarRocks: bucket and compaction pressure. The worst day is a table created with the wrong bucket count, discovered after it has terabytes in it. Too few buckets limit parallelism; too many create tiny tablets and metadata overhead on the frontend. Primary-key tables add memory pressure for the key index, so sizing a primary-key table by raw data volume alone underestimates its footprint. StarRocks 4.1’s range-based tablet splitting is aimed at precisely this problem, but it is new.
Shared-data cold reads. In any shared-data deployment, the first query touching cold data pays object-storage latency. A dashboard that looks fine in a warm test can have a bad first-morning experience. Warm caches deliberately, as Doris’s event-driven warmup allows, and measure cold latency separately.
Vendor benchmark gravity. Every number in this post attributed to a vendor was produced on that vendor’s chosen workload and hardware. ClickBench, TPC-H, TPC-DS and SSB are useful for regression tracking and poor for cross-engine purchasing. A representative test set of your twenty slowest real queries, run against realistic data volume and concurrency, outweighs any published chart.
Experimental features in production. The ClickHouse Cascades optimizer is experimental in 26.8, StarRocks inverted indexing is beta in 4.1, and new minor lines have had early container problems. The rule we apply: new minor versions wait for the third patch release unless you need a specific fix.
Resharding and rebalancing. Moving a ClickHouse cluster from three shards to six is a project. In Doris and StarRocks, adding backends triggers automatic tablet rebalancing, but heavy rebalancing competes with query traffic. In all three, plan growth headroom rather than assuming elasticity.
Mixing workloads. Putting search, vector retrieval, and telemetry in one cluster saves a system but couples blast radius. A runaway vector build should not be able to slow an alarm dashboard. Use resource groups or workload isolation features, or separate clusters.
Practical Recommendations
Start from your workload shape, not from the engine.
If your data is mostly one wide, append-only event table and your questions are scans, filters, and aggregations, choose ClickHouse. It is the simplest to run, it compresses telemetry exceptionally well, and 26.8 LTS gives you a stable base. Handle the occasional join with dictionaries or denormalization. If you expect multi-join dashboards later, keep the schema join-friendly and revisit when the Cascades optimizer graduates from experimental.
If your product serves many users from a star schema with several joins, or your tables hold state that changes, choose StarRocks or Doris. Use colocation for the largest join, primary-key tables for mutable state, and a duplicate-model table for raw telemetry. Choose Doris when search and vectors are on the roadmap; choose StarRocks when concurrency and tenant skew dominate.
If you need both profiles, separating them is legitimate: ClickHouse for the telemetry firehose, an MPP engine for the serving layer that joins rolled-up facts to dimensions, with Iceberg as the contract between them.
Decision checklist
- Count the joins in your twenty most important queries. Three or more large-table joins per query points to Doris or StarRocks.
- Measure the fraction of rows updated after landing. Anything beyond a small percentage points to a primary-key engine.
- Define your concurrency target in queries per second and your P99 latency budget, then test at that load.
- Decide your Iceberg catalog and write path, and test it end to end on the candidate engine.
- Pick a version policy: ClickHouse 26.8 LTS, or the third patch of a new minor line for Doris and StarRocks.
- Run a two-week proof of concept on your own data, with cold-cache and warm-cache latency measured separately.
- Record the decision, the drivers, and the accepted consequences in your own ADR so the next team understands why.
Frequently Asked Questions
Which is faster, ClickHouse, Doris, or StarRocks?
It depends on query shape. For single-table scans and aggregations over wide event tables, all three are fast enough that the winner varies by schema and hardware, and published benchmarks are vendor-run. For multi-table joins, StarRocks and Doris have the more mature cost-based distributed planners, while ClickHouse is catching up, with a Cascades optimizer experimental in 26.8. Test your own twenty slowest queries at realistic scale and concurrency rather than trusting a chart.
Is StarRocks a fork of Apache Doris?
Yes, in origin. StarRocks started from a Doris code lineage and then rebuilt major parts of the query engine, including the optimizer and a pipeline execution model, so the two have diverged substantially. They still share a recognizable architecture of frontends and backends, similar table models, and MySQL-protocol compatibility. Doris is governed by the Apache Software Foundation, while StarRocks is released under Apache License 2.0 and hosted under the Linux Foundation. Treat them as cousins, not as one product.
Can ClickHouse handle joins now?
Better than before, with caveats. ClickHouse 26.8 added statistics collection on insert for small tables, an IEJoin algorithm for inequality joins, a parallel full-sorting-merge join, and an experimental Cascades optimizer for distributed join strategy. Simple joins with a small right side have long worked well. Star schemas with several large dimension tables still reward denormalization, dictionaries, or an MPP engine until the distributed planner leaves experimental status and proves itself on your workload.
Which engine is best for IoT and digital twin telemetry?
For a high-volume, append-only telemetry firehose with scans and rollups, ClickHouse is often the simplest and most compact choice, partly because codecs such as DoubleDelta and Gorilla compress sensor data well. If your twin serves many users from joins between telemetry, asset hierarchies and mutable device state, StarRocks or Doris fit better. Many teams use both, with Iceberg between them. Match the engine to query shape rather than to the industry label.
Do I need storage-compute separation?
Only if your load is spiky or your cold data dominates cost. Shared-data mode on object storage lets you scale compute independently and pay object-storage prices for history, but cold reads are slower and cache management becomes your job. If your workload is steady and latency-critical, local disks in shared-nothing mode are simpler and faster. A middle path is to keep a small hot tier local and move older partitions to Iceberg on object storage.
Should I wait for a new minor version before upgrading?
For production, usually yes. Both StarRocks and Doris publish frequent patch releases, and the StarRocks 4.1 notes warn container users off v4.1.0 for a backend startup issue. A sensible policy is to adopt a new minor line at its third patch, or to stay on an LTS such as ClickHouse 26.8. Always read the release notes for behavior changes, such as StarRocks 4.1.6 making online table optimization opt-in.
Further Reading
- Industrial IoT time-series platform architecture for the ingestion layer that feeds these engines.
- ClickHouse 26.8 LTS vs 26.3 for the release-to-release changes in ClickHouse.
- ClickHouse vs Druid vs Pinot and Apache Pinot vs Druid for append-only stream engines.
- DuckDB vs ClickHouse embedded analytics for single-node alternatives.
- Postgres 18 vs TimescaleDB vs ClickHouse for IoT for smaller-scale time-series choices.
- Apache Iceberg v3 spec features and upgrade for the table format that sits between hot and cold tiers.
- NATS JetStream vs Kafka for edge IIoT telemetry for the log in front of the database.
- External: Apache Doris documentation and StarRocks documentation.
References
- ClickHouse, “ClickHouse Release 26.8” (blog and release call presentation), September 2026: https://clickhouse.com/blog/clickhouse-release-26-08 and https://presentations.clickhouse.com/2026-release-26.8/
- Apache Doris, release notes index (version and date list): https://doris.apache.org/releases/core/
- Apache Doris 4.0.7 release notes: https://doris.apache.org/releases/v4.0/release-4.0.7/
- Apache Doris 4.0 announcement: https://dev.to/apachedoris/apache-doris-40-one-engine-for-analytics-full-text-search-and-vector-search-k8n
- VeloDB, “Apache Doris 4.1: unified storage and retrieval for AI and search”: https://www.velodb.io/blog/apache-doris-4-1-unified-storage-and-retrieval-for-ai-and-search
- StarRocks 4.0 release notes: https://docs.starrocks.io/releasenotes/release-4.0/ and announcement https://www.starrocks.io/blog/starrocks-4.0-now-available
- StarRocks 4.1 release notes: https://docs.starrocks.io/releasenotes/release-4.1/ and announcement https://www.starrocks.io/blog/starrocks-4.1-now-available-built-for-production-designed-to-simplify
- StarRocks, “StarRocks Is Now Under Apache License 2.0”: https://www.starrocks.io/blog/starrocks-is-now-under-apache-license-2.0
- ClickHouse documentation, MergeTree and join algorithms: https://clickhouse.com/docs
By Riju – about
