Databricks 用 2026 基准测试论证湖仓比数据仓库更适合 AI 时代
The lakehouse is a better data warehouse: 2026 benchmarks and proof
Databricks 发布 2026 基准测试与客户案例,论证湖仓(Lakehouse)在 SQL 性能上已追平数据仓库:其生产 SQL 工作负载较 2022 年基线快 77%,Performance Index 约提升 4 倍。
The warehouse was the right answer for 20 years
The data warehouse earned its place. For two decades, it was the most reliable way to store structured data, enforce schemas and serve SQL queries to dashboards and reports. If your workload was "analysts writing SQL against clean tables," the warehouse was hard to beat.
That workload still exists. But it's no longer the only workload that matters, and it's no longer the workload that determines competitive advantage. In 2026, the organizations pulling ahead are the ones whose data platform serves BI analysts, data engineers, ML models, streaming pipelines and AI agents from the same governed foundation. The warehouse was built for one of those consumers. The lakehouse was built for all of them.
This post makes the case with benchmarks, customer evidence and architectural analysis. We'll be direct about where the warehouse still has strengths, because credibility matters more than cheerleading.
How lakehouse SQL performance caught up
The most common objection to the lakehouse has always been SQL performance. "Sure, it's flexible, but can it match my warehouse for dashboards?" In 2026, the answer is yes.
The Databricks Lakehouse trajectory
Databricks reported that its production SQL workload mix was 77% faster in late 2024 than its 2022 baseline, equivalent to roughly a 4x improvement in the Databricks Performance Index. That index is derived from billions of production queries across BI, ETL and data exploration workloads, not a single synthetic benchmark.
The improvements came from compounding gains across the Photon vectorized execution engine, query planning, caching, predictive optimization, join algorithms and concurrency management. Specific workload categories improved approximately 14% for BI, 13% for data exploration and 9% for ETL over a recent five-month measurement period.
What customers actually measured
The strongest evidence isn't synthetic benchmarks, it's production results. Lumen Technologies migrated 133 TB of mission-critical telecom data from Cloudera/on-premises to Databricks Lakehouse and measured a 90% query performance improvement, with queries that took hours completing in minutes. Lumen also cut compute costs by 30-40% using Databricks Lakehouse Serverless. Trek Bicycle saw 80-90% faster retail analytics after moving to the lakehouse, going from one daily refresh to three. These aren't lab results. They're production workloads serving real business operations.
Real-time analytics: Lakehouse//RT
For workloads that need sub-second latency at high concurrency, Lakehouse//RT is a serverless SQL warehouse designed specifically for analytical reads. Databricks reports as low as 10 ms and up to 12,000 queries per second. It queries governed lakehouse tables directly, eliminating the need for a separate OLAP serving copy, CDC pipeline or synchronization layer.
This matters because the traditional architecture for real-time analytics required maintaining a separate serving database alongside the warehouse. The lakehouse collapses that into one governed layer.
What the warehouse still can't do
Matching SQL performance was the prerequisite. The real argument for the lakehouse is what it does beyond SQL.
1. Serve AI agents and ML models natively
Traditional warehouses were built for SQL users and BI tools. AI agents need governed access to structured tables, unstructured documents, embeddings and real-time signals from the same platform. Gartner describes this shift as a movement from the warehouse toward a broader "analytical control plane" that combines governed data, semantic layers, AI agents and execution capabilities.
Deloitte reported in 2026 that only about one in five companies had a mature governance model for autonomous AI agents. We believe the path forward requires platforms that govern data, models and agent behavior together, not warehouses that govern only tables
The lakehouse treats AI as a first-class workload. The same governed tables that power dashboards also power model training, feature engineering, retrieval-augmented generation and agent tool-calling. Unity Gateway extends governance to the AI runtime: controlling which models agents can call, logging prompts and responses, enforcing rate limits and governing MCP servers and tools.
2. Handle unstructured and semi-structured data
Warehouses excel at rows and columns. But enterprise knowledge lives in documents, emails, PDFs, images, logs and code. A 2026 Komprise survey identified data classification and tagging as the leading challenge in preparing unstructured data for AI, while McKinsey's research on AI data readiness highlights the importance of connecting unstructured content with structured enterprise data.
The lakehouse stores all data types in open formats on cloud object storage. Structured tables, semi-structured JSON, unstructured documents and vector embeddings coexist under the same governance model. You don't need a separate system for each data type.
3. Eliminate data copying between systems
This is where the TCO argument gets real. Most warehouse architectures require data to flow from a lake to the warehouse, then get copied again into ML environments, feature stores or departmental extracts. Every copy costs storage, compute and governance overhead.
The lakehouse eliminates this by serving all workloads from the same governed data. SQL queries, Spark transformations, streaming pipelines, Python notebooks and ML training jobs all read from the same Delta Lake or Apache Iceberg™ tables. No copying, no synchronization, no lineage breaks.
4. Deliver streaming and batch on one platform
Warehouses are batch-first. Data arrives on a schedule, gets transformed and becomes queryable hours later. The lakehouse supports streaming ingestion alongside batch, with data becoming queryable in seconds rather than hours. For use cases like fraud detection, operational monitoring and real-time personalization, it's a requirement and not a nice-to-have.
5. Run on open formats you control
Warehouses that tightly couple storage, compute and proprietary data formats create switching costs that compound over time. The lakehouse runs on Delta Lake and Apache Iceberg™, open table formats that any compatible engine can read. Unity Catalog provides an open, multi-engine catalog layer with Apache 2.0 licensing, and UniForm generates Iceberg-compatible metadata from Delta tables so external engines can read the same data without copying it.
The proof: what organizations actually measured
Customer stories are where the argument moves from theory to evidence. Here's what organizations measured after migrating from warehouses to the lakehouse.
Cost reductions
Organization | Migration | Cost result |
Legacy warehouse to Databricks Lakehouse | 60% warehouse cost reduction, 70% lower ETL costs | |
Oracle + Azure Synapse + custom warehouses to Databricks | 30% reduction in data-landscape costs | |
SQL warehouses to Databricks Serverless | 29% overall cost reduction, 33% lower idle-compute costs | |
Amazon Redshift to Databricks Lakehouse | 25-30% lower TCO |
The pattern is consistent: 25-75% cost reductions, driven not just by cheaper queries but by eliminating redundant systems, duplicate data copies and multi-platform governance overhead.
Performance improvements
Organization | Migration | Performance result |
Cloudera/on-premises to Azure Databricks Lakehouse | 90% query performance improvement (hours to minutes), 133 TB migrated | |
Legacy warehouse to Databricks Lakehouse | 80-90% faster analytics, 3x daily refreshes (from 1x) | |
Snowflake to Databricks Lakehouse | 40% increase in development productivity |
Migration speed
Organization | Scale | Timeline |
1.5 PB across 44 business areas | 15 months (vs. original 2.5-year estimate) | |
Redshift migration | ~8 weeks | |
133 TB, mission-critical telecom systems | 9 months, on time and on budget |
Migration timelines have compressed dramatically. Lakebridge automates assessment, transpilation and reconciliation across Snowflake, Oracle, Teradata, Redshift, BigQuery and SQL Server. The Databricks Migration Agent uses AI agents to translate SQL scripts across dialects, converting up to 300 files per batch. These tools don't eliminate the need for human validation, but they reduce the manual effort that made migrations prohibitively slow.
Where the warehouse still wins (and why it matters less)
Here's where the warehouse retains advantages:
- Predictable out-of-the-box SQL performance. A well-tuned warehouse requires less data engineering effort to deliver fast dashboard queries. The lakehouse's performance is more sensitive to file sizes, partition strategy, clustering, statistics and compaction. A poorly maintained lakehouse can be dramatically slower than a warehouse despite identical hardware.
- Simpler administration for SQL-only workloads. If 90%+ of your compute is structured SQL dashboards and your team is entirely SQL-focused, a warehouse may require less operational overhead.
- Mature high-concurrency BI serving. Snowflake's multi-cluster warehouses and Interactive Warehouses are purpose-built for thousands of concurrent short queries. Lakehouse SQL engines are competitive but may require more configuration to match this specific pattern.
But these advantages matter less in 2026: the percentage of enterprise data workloads that are "SQL-only BI" is shrinking. Forrester's Q3 2026 Wave evaluated 14 leading lakehouse vendors. Gartner published its first Market Guide for Data Lakehouse Platforms in 2025 and followed with a Lakehouse Reference Architecture in 2026. Both firms now treat the lakehouse as a mainstream enterprise data-platform category. This tells us the warehouse's strengths are real but increasingly narrow as AI, streaming, and unstructured data become core enterprise requirements.
The architecture that's winning
The organizations reporting the strongest results aren't running a lakehouse instead of a warehouse. They're running a lakehouse that is a better warehouse, plus everything else.
The architecture looks like this:
- One governed data foundation built on open table formats (Delta Lake, Apache Iceberg™) on cloud object storage, governed by Unity Catalog.
- Multiple workload engines running over the same data: Databricks Lakehouse for SQL analytics, Lakehouse//RT for sub-second serving, Apache Spark™ Declarative Pipelines for ETL, Lakeflow Jobs for orchestration, Agent Bricks for AI agents and Lakebase for operational/transactional workloads.
- Unified governance spanning data, models, agents, tools and AI traffic through Unity Catalog and Unity Gateway.
- Open interoperability through UniForm (Delta tables readable as Iceberg), Unity Catalog (open multi-engine catalog) and standard JDBC/ODBC connections for Power BI, Tableau, Looker and other BI tools.
What to do if you're still on a warehouse
You don't need to migrate everything tomorrow. But you should start measuring the gap.
- Quantify your hidden costs. Add up what you spend on the warehouse, the separate ML platform, the streaming infrastructure, the data lake, the governance tools and the data copies between them. That's your real TCO, not just the warehouse bill.
- Run a side-by-side proof of concept. Pick one high-value domain. Run the same queries on your warehouse and on Databricks Lakehouse. Measure p50, p95 and p99 latency at realistic concurrency. Measure cost per 1,000 queries. Then add the ML and streaming workloads the warehouse can't serve and measure the total platform cost.
- Start with new workloads. If migration feels risky, build your next AI, streaming or data engineering project on the lakehouse. Let the warehouse handle existing BI while the lakehouse proves itself on new use cases.
- Use Lakebridge to assess migration complexity. The assessment phase is free and gives you a concrete picture of what migration involves for your specific SQL estate.
- Set explicit migration gates. Don't migrate until the lakehouse matches or beats the warehouse on query performance, data reconciliation, availability and security for each domain.
The warehouse was the right answer for a long time. The lakehouse is the right answer now, not because the warehouse broke, but because the question changed.
Get started:
Frequently asked questions
Is the lakehouse actually faster than a data warehouse for SQL queries? For most workloads, the lakehouse now matches or exceeds the performance of warehouses for SQL. Databricks Lakehouse has improved 77% since 2022, and independent TPC-DS testing shows sub-second p50 latency on 1-TB workloads at 74% lower cost than classic compute. The warehouse may still edge ahead on highly concurrent short queries with warm caches, but the gap is narrow and closing.
How much can I save by migrating from a warehouse to a lakehouse? Published customer results range from 25% to 75% cost reduction. AXA Japan reduced warehouse costs by 60%. The savings come from eliminating redundant systems and data copies, not just cheaper queries.
What about my existing BI tools like Power BI and Tableau? They work natively with Databricks Lakehouse through standard JDBC/ODBC connections. Your existing dashboards and reports connect to lakehouse tables without rebuilding. AI/BI Dashboards and Genie Agents add natural-language analytics on top.
How long does a warehouse-to-lakehouse migration take? It depends on complexity, but timelines have compressed. Vivriti Capital migrated from Redshift in ~8 weeks. IndusInd Bank migrated 1.5 PB across 44 business areas in 15 months. Lakebridge and the Agentic Code Converter automate assessment, SQL translation and data reconciliation to accelerate the process.
Can the lakehouse handle real-time analytics? Yes. Lakehouse//RT delivers sub-second analytical queries at high concurrency directly on governed lakehouse tables. Databricks reports sub-100-ms latency at 12,000 queries per second. For streaming ingestion, Zerobus and Apache Spark™ Declarative Pipelines provide sub-second data freshness.
来源:Databricks:Blog · databricks.com