LinkedIn has consolidated three systems for finding and analysing metric metadata into one ClickHouse index, according to an account published by ClickHouse on 1 October. LinkedIn engineer Jacob Zelek estimates that the migration reduced memory use to roughly one-fifth and compute to about two-thirds of the previous stack, while supporting queries it could not previously answer. The index serves more than 150,000 queries a minute across over 13 billion metrics, with a reported average latency of 68 milliseconds. That is a useful result for anyone whose monitoring infrastructure has become something that itself needs monitoring.
ClickHouse Watch analysis
What happened
LinkedIn’s older metric-discovery service relied on three systems: two custom in-memory stores and a separate search service. Together, they had reached their scaling limits by 2025, yet still could not handle some of the analytical questions engineers wanted to ask.
The replacement preserves the gateway that existing tools already used. Instead of distributing requests across those three systems, it translates them into database queries against the consolidated index. Downstream tools did not need to change, making the migration transparent to their users.
That matters because LinkedIn has millions of graphs and alerts built around legacy metric names. These names combine several dimensions into a single string. The new index exposes those dimensions separately, letting engineers filter directly rather than search the entire name with a regular expression.
LinkedIn’s metrics also have roughly 30% daily turnover. A continuous change feed keeps the index current, while a fresh daily table is built, brought up to date and swapped into service. The reported deployment holds the metadata on a single shard, with estimated headroom to double its metric volume.
Read the LinkedIn observability account.
Why it matters
This is metric metadata discovery and analysis, not a claim that LinkedIn stores every underlying measurement in this index. Engineers use it to answer questions such as which hosts emit a metric, which metrics belong to a service and how metric counts are growing.
Consolidation brings two benefits here: fewer systems to operate and a broader set of questions the same data can answer. The 68-millisecond figure is an average across query types, including heavier analysis, not a guarantee about the slowest requests.
The team also had operational experience to build on. LinkedIn had already deployed ClickHouse for distributed tracing across three data centres, handling around 800 billion spans a day. Familiarity with the database helped make this second project easier to adopt.
Our read
The strongest lesson is the migration design, not simply the impressive number of zeroes. LinkedIn kept the interface its tools depended on, replaced the machinery behind it and made individual metric dimensions easier to query. A platform migration that does not require every team to rearrange its furniture deserves some appreciation.
For infrastructure teams, the useful starting point is to map which discovery queries your current systems cannot answer, then test whether a consolidated index can support both those questions and existing traffic. Measure resource use and slow-query behaviour separately. LinkedIn’s reported savings are a promising case study, not a budget forecast for everyone else.
What to watch
- Whether rewriting legacy pattern searches into direct dimension filters delivers further resource savings.
- How query latency and capacity change as metric volume grows towards the estimated doubling headroom.
- Whether LinkedIn can retire more legacy naming dependencies without disrupting existing graphs and alerts.
Discussion spark: Should observability teams consolidate metric discovery into one analytical database, or keep specialised systems when workload isolation matters more than operating simplicity?
Sources and evidence
- Source update (1 October 2026, 15:50 UTC)
not affiliated with or endorsed by ClickHouse