On August 3, 2026, Andy Pavlo announced that he's leaving Carnegie Mellon University to join ClickHouse and establish ClickHouse Labs, the company's first formal research organization. Its proposed work includes database design for autonomous agents, deeper analysis of existing engineering, and new approaches to hardware and query execution.

The hire comes as ClickHouse reports rapid growth among AI companies. Those customers need to analyze large volumes of operational data quickly, often while applications are still running. A dedicated research lab gives ClickHouse room to investigate problems that extend beyond the next product release, including how databases should behave when the software querying them is an AI agent.

Pavlo's background fits the research agenda

Pavlo has been an Associate Professor of Databaseology at CMU since 2013. His awards include the 2014 ACM SIGMOD Jim Gray Dissertation Award, a 2018 Sloan Research Fellowship, and a 2019 NSF CAREER Award. This year, he received the 2026 IEEE TCDE Ramez Elmasri Outstanding Database Education Award, recognizing a decade of freely available course materials and lectures.

Those teaching materials have given engineers a way to understand database internals beyond the interfaces their applications use. His research and software projects also connect directly to the questions ClickHouse Labs plans to explore:

  • NoisePage was a self-driving database management system, or DBMS, project. It used machine learning to forecast workloads, model database behavior, and generate tuning actions without a database administrator directing each change.
  • OtterTune, the database tuning startup he co-founded in 2020, brought related ideas into a commercial product. It used query performance data to guide automated tuning of database settings and indexes, aiming to perform work that otherwise required an experienced operator.
  • BusTub is the educational DBMS used in CMU's 15-445 course. It gives students practical experience with components such as buffer pool managers, B+ trees, and concurrency control.

Pavlo has also maintained DBDB.io, a community catalog of database systems from industry and academia. He's followed ClickHouse since its open-source release in 2016 and assigned its technical articles to his CMU students. His familiarity with the system gives the appointment more substance than a prominent academic name on an organization chart.

The business behind the lab

In January 2026, ClickHouse closed a $400 million Series D led by Dragoneer, at a reported valuation of around $15 billion. It also acquired Langfuse, the open-source platform for observing and evaluating large language model applications.

At its Open House 2026 conference, ClickHouse reported $250 million in annual recurring revenue, triple the year-earlier figure, and more than 4,000 ClickHouse Cloud customers. Anthropic, Meta, Cursor, Tesla, OpenAI, Weights & Biases, and Vercel were among the customers that appeared on stage to discuss their use of the system, alongside more than a dozen others.

According to ClickHouse, AI companies are its fastest-growing customer segment. There is a practical reason for that demand. Large language models and their supporting applications produce inference logs, token counts, latency traces, embedding metadata, and user interaction records. These are structured operational records that teams need to query, sometimes with sub-second response times and many requests running at once.

ClickHouse's design is well suited to this kind of analytics. Columnar storage groups values by column, reducing unnecessary reads when a query needs only a few fields. Vectorized execution processes batches of values rather than handling each row separately. Compression reduces storage requirements and the amount of data that has to move through the system. Together, these choices support fast analytical queries over large datasets.

At Open House, CEO Aaron Katz said AI workloads demand the performance and cost efficiency ClickHouse was built to provide. That's the company's assessment, but it also describes the opportunity behind the research investment. An analytics database serving human-authored dashboards faces different demands from one supporting AI systems that may query millions of rows per second while deciding what to do next.

Designing databases for autonomous agents

ClickHouse Labs' stated ambition is to produce research with an impact comparable to IBM Research or Microsoft Research. That's a high target. The more useful way to assess the lab is through the specific problems it intends to work on.

Database design for autonomous agents is the most interesting direction on that list. Agents can generate queries dynamically, work with an incomplete understanding of the available schema, and run many small queries in tight loops. Their access patterns can differ from both a human analyst writing SQL and conventional application code following a predictable sequence.

That creates research questions beyond making an individual query faster. A database might predict an agent's next query and prepare results in advance. It might adjust execution plans based on observed agent behavior or surface relevant data that the agent hasn't requested yet. These are possible design directions, not capabilities established by the lab's announcement.

Pavlo's work on NoisePage and OtterTune is relevant here. Both projects examined how a database could use observations about its workload to make operational decisions. Supporting agents adds another layer: the system may need to accommodate software that is still discovering what data exists and how to use it.

The argument for this research is that agent access remains an underexplored database design problem. Optimizing for that behavior could lead to different choices from those made for a warehouse primarily serving human-authored SQL.

Testing existing work more thoroughly

A second direction is accelerating the validation of engineering ClickHouse has already done. Research groups can examine ideas that have been implemented but haven't yet received extensive formal analysis or comparative testing.

ClickHouse has shipped substantial engineering work over the past three years. A lab can apply formal methods, disciplined benchmarks, and comparisons with alternative approaches to that work. These methods serve different purposes. Formal analysis can examine whether a design has particular correctness properties, while benchmarking can show where its performance holds up and where it falls short.

A product engineering team under delivery pressure may have limited time for that depth of investigation. Dedicated researchers can test assumptions and identify limits that are difficult to expose through routine development. This work is less visible than a new feature, but it can help establish which designs are ready for broader production use.

Preparing for different hardware

The third direction covers new hardware, algorithms, and execution strategies. Technologies such as CXL memory, which can change how systems attach and share memory, and disaggregated storage, which separates storage resources from compute, create different constraints for database engines. AI accelerators increasingly designed alongside software add more possibilities.

These developments can change assumptions about memory bandwidth, cache behavior, and the cost of moving data. An execution strategy that works well on one hardware layout may perform differently on another. Investigating those differences early could help ClickHouse adapt its engineering roadmap before new hardware becomes common in customer deployments.

A research lab provides a place to explore those questions without requiring every experiment to become an immediate product feature. Some ideas may improve the existing engine. Others may show that a promising technology doesn't offer enough benefit for the workloads ClickHouse serves.

What the commitment signals

Establishing a research organization suggests that ClickHouse intends to invest beyond its current growth cycle. Funding provides the resources for that work; a lab creates a structure for pursuing questions whose commercial value may take years to become clear. It doesn't guarantee an IBM Research-scale result, but it is a meaningful long-term commitment.

The timing looks sensible. AI companies need fast analytical infrastructure, while enterprises are considering alternatives to expensive, tightly coupled cloud data warehouses. ClickHouse has the resources to fund research while competition around these workloads is still developing. If the analytics database market changes substantially over the next five years, the lab could give the company more influence over that direction.

The appointment gives engineers and architects a reason to examine ClickHouse's roadmap more closely, especially for AI systems that access structured data at scale. Following the lab's research will mean looking at how it characterizes agent workloads and what it finds about existing database designs, then tracking which results reach the product.