Databricks 为云数据仓库成本设置预算与告警
Set Budgets and Alerts for Cloud Data Warehouse Costs
Databricks 在 Data & AI Platform 上推出主动式成本管理方案,通过成本归因、预算、告警、仪表盘和优化五大支柱,让团队在花钱前就设好护栏。
Every analytics team has been here: a month-end invoice arrives, and a single SQL warehouse has blown through its budget, maybe a compute left running overnight, maybe an ad hoc one that auto-scaled to serve 80 Tableau users during a quarterly review. The cost is real, but the frustration runs deeper: nobody knew until it was too late.
Proactive cost management flips that model. Instead of forensic bill analysis after the fact, you set guardrails before a dollar is spent - budgets are tied to warehouses, teams, or BI workloads. Alerts that fire at the thresholds you choose, and dashboards that surface usage trends in real time, all on the Databricks Data & AI Platform. If you're mid-migration from an on-premises system or another cloud data warehouse, now is the moment to put that governance in place.

Why does reactive cost management fail?
Data warehousing workloads present unique cost challenges. SQL warehouses serve interactive, human-driven workloads (BI dashboards, ad hoc analysis, embedded analytics) where demand is unpredictable and bursty. A single executive all-hands that triggers 200 simultaneous dashboard refreshes can spike costs in minutes. Most organizations start their FinOps journey the same way: someone downloads last month's usage CSV, opens a pivot table, and starts asking questions. This approach breaks down for four reasons:
- You can't see who's spending. Usage rolls up to the platform, not to the team or query that caused it, so that no one can be held to a number.
- No one owns a budget. Without spending targets tied to a warehouse or team, the analyst running SELECT * on a 2 TB table never feels the cost. The platform team quietly absorbs it.
- You find out too late. By the time last month's Excel reports surface the overspend, it's two billing cycles old and the decision that caused it is invisible.
- Manual review doesn't scale. Five warehouses may pass a monthly pivot-table review. However, once scaled to dozens, supporting hundreds of analysts across Tableau, Power BI, and Looker becomes unmanageable.
Proactive cost management — attribution, budgets, alerts, dashboards, and optimization — addresses all five.
Start with serverless: the first cost decision.
Before setting up budgets, choose the right warehouse type. SQL Serverless warehouses are Databricks’ recommended option for most workloads. They start in seconds, scale down the moment queries finish, and let Intelligent Workload Management allocate compute for you, so you pay only for actual query execution, not idle readiness.
If you're migrating from a legacy cloud data warehouse, this is a major shift. Most legacy platforms charge for always-on compute or require manual scaling decisions. Serverless removes that entire class of cost mistakes.
Lumen Technologies saw this firsthand: Migrating two essential telecom systems from an on-premises data warehouse to SQL Serverless reduced compute costs by 30 -40% and increased query speed by 90%. The system now automatically scales to handle approximately 7 GB every 10 minutes, with Serverless removing the "always-on tax" typically associated with bursty, real-time workloads.
The five pillars of proactive cost management
The framework consists of five pillars. The first four make your spend visible; the fifth is what actually brings it down:
- Cost Attribution. Tag everything so you know who spent what on which warehouse.
- Budgets. Define monthly spending targets at the account, workspace, or team level.
- Alerts. Get notified the moment the spend approaches or exceeds a threshold.
- Dashboards. Visualize warehouse trends and make cost a first-class operational metric.
- Optimizations. Reduce what the work costs in the first place.

Pillar 1: Cost attribution at two levels
You can't manage what you can't attribute, and with warehouses that means two levels: the warehouse itself and the individual query. One tells you which warehouse spent the money; the other tells you who inside it did.
Level 1: Tagging the warehouse
SQL warehouses support custom tags as key:value pairs that propagate to system.billing.usage, linking every DBU consumed to a team, project, or cost centre. Tagging works the same way for every warehouse type — Classic, Pro, and Serverless are all tagged at the warehouse level, so you learn one pattern and apply it everywhere.
Set tags in the UI under SQL Warehouses > your warehouse > Edit > Tags, or create the warehouse through the REST API:
Tagging through the REST API
Apply at least three dimensions: team (who owns it), cost center (who pays), and use case (BI serving, ad hoc, scheduled refresh).
There isn't a native policy today that enforces warehouse tags, so anyone can create an untagged one. Build enforcement into your provisioning process instead — create warehouses with Terraform, Declarative Automation Bundles, or the REST API with tags in the definition, and treat the UI as the exception. The attribution-gap query later in this post catches whatever slips through.
Level 2: Tagging the query
Warehouse tags tell you a warehouse cost $4,000 last month. They don't tell you that 60% of it came from a single Power BI dashboard refreshing every 15 minutes. That gap matters because warehouses are shared — one BI warehouse serves dozens of dashboards, several dbt models, and a stream of ad hoc queries.

Query tags (Public Preview) close it by attaching business context to individual SQL statements:
The tags land in the query_tags column of system.query.history (Public Preview), alongside executed_by and statement_id, so you can group spend by dashboard, model, or cost centre rather than by warehouse. Some tools set these for you: starting with dbt-databricks 1.11.0, model queries are automatically tagged with the dbt model name, and Power BI passes workspace and dataset identifiers through the ADBC driver.
Two caveats: query tags apply to SQL warehouse queries only, and because they live in query history rather than billing data, you join the two yourself to get per-query cost.
Before writing any SQL, check the Cost page in Governance Hub (Beta) — an account-level view of spend, cost drivers, budgets, and tagging coverage, and the quickest way to see how much of your usage is attributable at all. An account admin enables it from the account console's Previews page.

Pillar 2: Setting budgets in the account console
Creating a budget
In the Account Console, go to Usage > Budgets, click Add budget, and configure:
- Name: Descriptive (e.g., "Analytics SQL Warehouses, Monthly")
- Amount: Monthly target in USD
- Scope: Filter by workspaces and/or tags (e.g.,Team:Analytics)
- Email notifications: Recipients notified when spend reaches the budget amount
Example budget configurations
Budget Name | Amount | Scope (Tags) | Alert Recipients |
|---|---|---|---|
BI Serving + Production | $8,000/mo |
| platform-team@company.com |
Analytics, Ad-Hoc | $3,000/mo |
| analytics-mgr@company.com |
Marketing Analytics | $2,500/mo |
| mkt-data-lead@company.com |
Account-Wide Safety Net | $25,000/mo | (all SQL usage) | cto@co.com, finops@company.com |
The key pattern is layered budgets: team-level budgets for accountability, plus an account-wide budget as a safety net that catches anything that slips through.
Keep in mind that budgets are a monitoring mechanism, not a hard cap. They don't stop usage or prevent charges, so your bill can still exceed the amount. The goal is awareness and rapid response, not shutoffs that could break a production dashboard mid-refresh.
GetYourGuide tested this directly: consolidating all their Looker workloads onto SQL Serverless lowered BI serving costs by ~20% and sped queries 35%, even though the team had expected Classic clusters to be cheaper. Warehouse-type choice is a budgeting decision worth re-validating, not a one-time call.
Reading your burn rate
Click any budget to see current spend vs. target, remaining budget, and a day-by-day burn-rate visualisation. For data warehousing, this chart is especially revealing:
- Linear daily burn means stable BI serving costs.
- Sawtooth with Monday/Friday spikes indicates ad hoc exploration clustering around workdays.
- A sudden step-up mid-month means a new dashboard was deployed without a cost review.
- Gradual upward drift means organic adoption is outpacing your budget assumptions.
Pillar 3: Alerts, the nervous system of cost governance
Budgets tell you where you stand. Alerts tell you when to act.
Budget email alerts
When you create a budget, add email recipients who are notified when spending exceeds the budget amount. Zero code is needed.
For threshold-based alerting, create Databricks SQL Alerts that query `system.billing.usage` directly. Here are the patterns that matter most:
Alert when the daily SQL warehouse spend exceeds a threshold
Pacing alert: projected month-end spend will exceed budget
This is the pattern teams get the most value from. Instead of alerting after the breach, it projects month-end spend and fires before that projection crosses your budget while you still have time to act:
Schedule this every four hours and route to Slack or email. It gives teams days to react (right-size a warehouse, optimize an expensive query, defer a batch) before the budget is breached.
Detect hidden costs from materialized view refreshes
Materialized view refresh costs are the ones teams miss most often, because they bill as serverless pipeline usage rather than warehouse usage. A view set to refresh every 15 minutes against a source that updates twice a day spends the difference for no benefit. Match the refresh schedule to how often the underlying data actually changes, then watch the pipeline spend directly:
Pillar 4: Dashboards and continuous monitoring
Alerts are point-in-time triggers. Dashboards give continuous context. Databricks provides pre-built usage dashboards that account admins can import into any Unity Catalog-enabled workspace. These cover spend trends, top-N analysis, tag filtering, and per-warehouse drill-downs. For custom monitoring, two queries are particularly useful:
SQL warehouse cost by team
Unattributed SQL usage (attribution gap)
Untagged SQL usage is your attribution gap, warehouse spend that can't be traced to any team. Drive it to zero.
Raiffeisen Bank International built its own cost-monitoring and forecasting platform directly on Databricks system tables, with per-warehouse, per-user, and per-workload visibility. The payoff was cultural as much as technical: 3-4x faster workloads and, more importantly, when teams can see their own spend and are expected to explain it, behaviour changes without mandates.
Pillar 5: Optimization, the pillar that actually changes the bill
The first four pillars are about seeing your spending. This one is about lowering it, and it's the one most teams skip. Start with the two changes that improve cost and performance at the same time.
Turn on predictive optimization for Unity Catalog managed tables. Databricks runs OPTIMIZE, VACUUM, and ANALYZE for you based on how each table is actually queried, and with CLUSTER BY AUTO it adjusts clustering keys as query patterns change. Less data scanned means faster queries and fewer DBUs, with no maintenance schedule to keep. It's on by default for accounts created on or after November 11, 2024, and reaching older accounts through 2026.
Default to serverless warehouses for the startup and idle economics described earlier — for most teams that's the bulk of the savings, with no ongoing effort.
From there, a few settings are worth tuning:
- Tighten auto-stop. Pro and Classic default to 45 minutes idle before stopping; serverless defaults to 10, and you can drop it to 5 in the UI or 1 via the API. Every idle minute on a BI warehouse after the last dashboard closes is waste.
- Set a statement timeout (Beta — enable the preview, then set it per warehouse via the API). One runaway query shouldn't burn a weekend of compute; keep it short on BI warehouses and longer on ETL.
- Designate a default warehouse (Admin Settings > Compute) so a quick look at ten rows doesn't wake an ETL-sized cluster, and start Small or Medium rather than over-provisioning.
- Size up or out on purpose. Scaling up — larger sizes, now to 5X-Large (Public Preview) — makes a single heavy query faster; scaling out (a higher max cluster count) serves more concurrent users. Reaching for the wrong one is a common and expensive mistake.
The data layer matters too: OPTIMIZE, VACUUM, and liquid clustering each cut the compute a query needs, and that compounds across thousands of BI queries a day.
Underneath all of it, most cost problems are query problems. A missing join key or a SELECT * against a wide table costs the same no matter how the warehouse is tuned — and the attribution from Pillar 1 is what tells you which queries and dashboards to go fix.
Key takeaways
1. Bake serverless and attribution into the platform from day one. Default to SQL Serverless warehouses for instant startup, autoscaling, and no idle time cost, and enforce custom tags from the start, because unattributed usage is invisible and invisible usage always grows.
2. Budget and alert proactively, at more than one level. Pair per-team budgets for accountability with an account-wide budget for anomalies, and alert on pace, not just thresholds, so you hear about an overrun with time to act rather than after the breach.
3. Make spending visible, and make it cultural. The most durable savings come from transparency, not restrictions, so put cost dashboards where teams work. As RBI found, when teams can see their own spend, behaviour changes without mandates.
Get started
This is your fastest path to get started, and be on top of the bills before it explodes:
- Create your first budget in the account console.
- Attribute usage by tagging your SQL warehouse and queries
- Monitor costs with system tables, query system.billing.usage for warehouse-specific insights
- SQL warehouse sizing and scaling behaviour, understand IWM and serverless economics
- Cost optimisation best practices, warehouse sizing, Photon, and Delta Lake optimisation
- From Chaos to Control: A Cost Maturity Journey with Databricks, the crawl-walk-run framework for FinOps maturity
Ready to take control of your cloud data warehouse costs? Set up a Databricks free trial account, create your first SQL Serverless warehouse, and configure a budget with email alerts in the Account Console. It takes five minutes and costs nothing.
For a deeper dive into cost observability, explore the billing system tables documentation and import the pre-built usage dashboard to start monitoring spend from day one.
来源:Databricks:Blog · databricks.com