Skip to content
Clarisights Knowledge Center home
InboxAsk a human

Connecting Databricks on Clarisights

Connect your Databricks lakehouse to Clarisights so first-party tables and views can be read alongside the rest of your marketing and analytics data. Clarisights queries Databricks read-only through a SQL Warehouse, which lets you blend your warehouse data with channel performance, attribution, and finance datasets.

At a glance

Connector typeData warehouse / lakehouse (read-only)
AuthenticationPersonal Access Token (PAT) on a SQL Warehouse
Permissions neededCAN USE on the SQL Warehouse and SELECT on the table or view
Object supportTables and views (Delta tables and Unity Catalog supported; Hive metastore also supported for legacy workspaces)
Refresh cadenceSet per pipeline
Limited rolloutYes — contact your Customer Success Manager to enable

Authentication options

Clarisights authenticates to Databricks using a Personal Access Token (PAT) tied to a SQL Warehouse in your workspace. The PAT is sent on every query and identifies the principal Clarisights runs as.

  • Strongly recommended: create the PAT against a service principal, not a personal user. Personal PATs break when the user leaves the org or rotates their credentials, which silently stalls the pipeline. A service-principal PAT keeps the connection stable across team changes.

  • If your workspace enforces IP allow-lists, you will also need to allow Clarisights' outbound IPs (shared during setup).

Setting up the connection

Step 1 — Create a service principal in Databricks

In your Databricks workspace, create a service principal that will own the connection to Clarisights. Using a service principal (instead of a personal account) keeps the integration stable when team members change.

[TODO: Screenshot needed — Databricks service principal creation screen]

Step 2 — Create a Personal Access Token for the service principal

Generate a PAT for the service principal you just created. Copy the token immediately — Databricks only shows it once. Set an expiry that matches your rotation policy.

⚠ Common error: Authentication failed when Clarisights validates the connection → caused by an invalid, expired, or revoked PAT → fix by generating a fresh PAT for the service principal and resharing it.

Step 3 — Grant permissions on the SQL Warehouse and the data

Grant the service principal:

  • CAN USE on the SQL Warehouse Clarisights will query through.

  • SELECT on each table or view that should be exposed to Clarisights.

Grant only the objects you want Clarisights to read — the connection cannot reach anything the service principal does not have SELECT on.

Step 4 — Share connection details with Clarisights

Send your CSM the values listed in Connection details exchange below. Clarisights will validate the credentials end-to-end before turning the pipeline on.

⚠ Common error: Warehouse not found or HTTP path invalid → caused by a mistyped SQL Warehouse HTTP path → fix by copying the path directly from SQL Warehouses → your warehouse → Connection details in Databricks (it looks like /sql/1.0/warehouses/<id>).

⚠ Common error: First query of the day takes 30 seconds to a couple of minutes → caused by SQL Warehouse cold start → fix by enabling auto-start on the warehouse, or keeping it warm during your pipeline window.

Connection details exchange

You provide to ClarisightsClarisights provides to you

Workspace URL (e.g. dbc-xxxxx.cloud.databricks.com)

Outbound IPs to add to your IP allow-list, if your workspace enforces one

SQL Warehouse HTTP path (e.g. /sql/1.0/warehouses/<id>)

Confirmation that credentials validated and the pipeline is scheduled

Personal Access Token (service principal recommended)

Catalog · schema · table or view name for each object to expose

What we read

Clarisights reads the tables and views you explicitly grant — nothing else. The connector supports:

  • Delta tables in your lakehouse.

  • Views, including views that join or transform underlying tables.

  • Unity Catalog three-part identifiers (catalog.schema.table).

  • Hive metastore two-part identifiers (schema.table) for legacy workspaces.

The query Clarisights runs is the SELECT configured during onboarding — we do not rewrite it, sample it, or add joins on the Databricks side.

Connector specifics

  • SQL Warehouse must be available when the pipeline runs. Either keep the warehouse running or enable auto-start. A cold warehouse can add 30 seconds to ~2 minutes to the first query of the day.

  • Use a service-principal PAT, not a personal PAT. Personal PATs are tied to a user account and stop working when that user leaves, rotates credentials, or loses workspace access. A service principal removes that fragility.

  • Identifier formats. Unity Catalog uses catalog.schema.table; the legacy Hive metastore uses schema.table. Both are supported — share the form your workspace uses when sending the table list.

Limitations & known constraints

  • Query timeouts. Long-running queries are subject to your SQL Warehouse's timeout settings. Heavy transformations should be materialized as a view or table and exposed to Clarisights via that object, not run as ad-hoc SQL on each pipeline run.

  • Cold-start lag. If the SQL Warehouse is stopped when Clarisights queries it, the first query has to wait for compute to come up. Enable auto-start to avoid pipeline failures during quiet hours.

  • Read-only. The connector only reads — Clarisights never writes back to your Databricks workspace.

Operating notes

  • Refresh cadence is set per pipeline during onboarding (e.g. hourly, daily). Tell your CSM the cadence you need; they will tune it against your warehouse capacity.

  • Rotate the PAT on schedule. When you rotate, share the new token with your CSM before the old one expires to avoid a gap in the pipeline.

  • Limited rollout. The Databricks connector is in limited rollout. Contact your Customer Success Manager to enable it for your workspace.

Need help?

When contacting support from the in-app messenger, please include:

  • The integration name and account ID (Integrations → Channel)

  • The exact error message or screenshot

  • The step where the issue occurred

  • When the issue started