> For the complete documentation index, see [llms.txt](https://docs.slingdata.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.slingdata.io/connections/database-connections/databricks.md).

# Databricks

Connect & Ingest data from / to a Databricks database

## Setup

The following credentials keys are accepted:

* `host` **(required)** -> The hostname of the Databricks workspace (e.g., `dbc-a1b2c3d4-e5f6.cloud.databricks.com`)
* `token` **(required)** -> The personal access token or password to access the instance
* `warehouse_id` **(required)** -> The SQL warehouse ID to connect to
* `http_path` (optional) -> The HTTP path for the connection (if not using warehouse\_id)
* `catalog` (optional) -> The initial catalog name to use in the session (default: `hive_metastore`)
* `schema` (optional) -> The initial schema name to use in the session (default: `default`)
* `port` (optional) -> The port number (default: `443`)
* `max_rows` (optional) -> Maximum number of rows fetched per request (default: `10000`)
* `internal_volume` (optional) -> Specifies a custom internal volume to use for bulk operations. If not provided, Sling will attempt to create a volume in the default schema named `SLING_SCHEMA.SLING_STAGING`.
* `table_format` (optional) -> The table format to use when Sling creates a table. Accepts `delta` or `iceberg` (default: `delta`). See [Table Format](#table-format) below.
* `copy_method` (optional) -> How Sling bulk-loads into Databricks. `stage` uses a Unity Catalog volume and `COPY INTO` (default). `aws` uses S3. `zerobus` streams Arrow RecordBatches into a Delta table (Beta). See [Zerobus](#zerobus) below.
* `timeout` (optional) -> Timeout in seconds for server query execution (no timeout by default)
* `user_agent_entry` (optional) -> Used to identify partners
* `ansi_mode` (optional) -> Boolean for ANSI SQL specification adherence (default: `false`)
* `timezone` (optional) -> Timezone setting (default: `UTC`)

### Using `sling conns`

Here are examples of setting a connection named `DATABRICKS`. We must provide the `type=databricks` property:

{% code overflow="wrap" %}

```bash
# Basic connection with warehouse
$ sling conns set DATABRICKS type=databricks host=<workspace-hostname> token=<access-token> warehouse_id=<warehouse-id>

# Connection with custom HTTP path
$ sling conns set DATABRICKS type=databricks host=<workspace-hostname> token=<access-token> http_path=<http-path>

# With catalog and schema
$ sling conns set DATABRICKS type=databricks host=<workspace-hostname> token=<access-token> warehouse_id=<warehouse-id> catalog=<catalog> schema=<schema>

# Or use url
$ sling conns set DATABRICKS url="databricks://token:<access-token>@<workspace-hostname>:443/sql/1.0/warehouses/<warehouse-id>?schema=<schema>"
```

{% endcode %}

### Environment Variable

See [here](https://docs.slingdata.io/connections/database-connections/pages/eAdVs2BHCgdr6RS8GoJC#dot-env-file-.env.sling) to learn more about the `.env.sling` file.

{% code overflow="wrap" %}

```bash
export DATABRICKS='databricks://token:<access-token>@<workspace-hostname>:443/sql/1.0/warehouses/<warehouse-id>?schema=<schema>'

# use JSON format
export DATABRICKS_CONN='{ "type": "databricks", "host": "<workspace-hostname>", "token": "<access-token>", "warehouse_id": "<warehouse-id>", "schema": "<schema>" }'

# use YAML format (with new lines)
export DATABRICKS='
type: databricks
host: <workspace-hostname>
token: <access-token>
warehouse_id: <warehouse-id>
schema: <schema>
'
```

{% endcode %}

### Sling Env File YAML

See [here](https://docs.slingdata.io/connections/database-connections/pages/eAdVs2BHCgdr6RS8GoJC#sling-env-file-env.yaml) to learn more about the sling `env.yaml` file.

```yaml
connections:
  DATABRICKS:
    type: databricks
    host: <workspace-hostname>
    token: <access-token>
    warehouse_id: <warehouse-id>
    schema: <schema>

  DATABRICKS_URL:
    url: "databricks://token:<access-token>@<workspace-hostname>:443/sql/1.0/warehouses/<warehouse-id>?catalog=<catalog>&schema=<schema>"\
```

## Table Format

Sling creates Delta tables by default. To create Managed Iceberg tables instead, set `table_format` to `iceberg` on the connection:

```yaml
connections:
  DATABRICKS:
    type: databricks
    host: <workspace-hostname>
    token: <access-token>
    warehouse_id: <warehouse-id>
    schema: <schema>
    table_format: iceberg
```

Or with `sling conns`:

{% code overflow="wrap" %}

```bash
$ sling conns set DATABRICKS type=databricks host=<workspace-hostname> token=<access-token> warehouse_id=<warehouse-id> table_format=iceberg
```

{% endcode %}

The format applies to the tables Sling creates, including the temporary tables used for `incremental` and `backfill` modes. All load modes and merge strategies work the same for both formats.

{% hint style="info" %}
Your workspace must have Managed Iceberg tables enabled in Unity Catalog. Sling does not change the format of a table that already exists.
{% endhint %}

To control the DDL further, use the [`table_ddl`](/concepts/replication/target-options.md#table-ddl) target option, which overrides `table_format`:

```yaml
target_options:
  table_ddl: |
    create table {object_name} ({col_types})
    using iceberg
    tblproperties ('key' = 'value')
```

## Zerobus

Set `copy_method: zerobus` to ingest with [Databricks Zerobus](https://docs.databricks.com/aws/en/ingestion/zerobus-ingest) instead of staging files. This is still the same `type: databricks` connection (there is no `type: zerobus`). The SQL warehouse is used for `DESCRIBE`, `CREATE TABLE`, `TRUNCATE`, and `MERGE`; Zerobus only streams rows into the Delta table.

This is Beta.

### Credentials

On top of `host`, `token`, and `warehouse_id`:

* `copy_method` **(required for Zerobus)** -> `zerobus`
* `client_id` **(required)** -> OAuth M2M service principal application ID
* `client_secret` **(required)** -> OAuth M2M client secret
* `zerobus_endpoint` (optional) -> Zerobus shard URL, **not** the workspace host. Example: `https://<workspace-id>.zerobus.<region>.cloud.databricks.com`. On Azure the suffix is `azuredatabricks.net`; on GCP it is `gcp.databricks.com`. If omitted, Sling tries to detect it from Unity Catalog.
* `batch_size` (optional) -> Rows per Arrow RecordBatch (default: `10000`)
* `ipc_compression` (optional) -> Arrow IPC compression: `none` (default), `lz4`, or `zstd`
* `max_inflight_batches` (optional) -> In-flight batches waiting for acknowledgment (default: `1000`)

Create a service principal and secret in the workspace, then grant it access to the target catalog/schema:

```sql
GRANT USE CATALOG ON CATALOG <catalog> TO `<client_id>`;
GRANT USE SCHEMA ON SCHEMA <catalog>.<schema> TO `<client_id>`;
GRANT SELECT, MODIFY, CREATE TABLE ON SCHEMA <catalog>.<schema> TO `<client_id>`;
```

See [Authorize service principal access to Databricks with OAuth](https://docs.databricks.com/aws/en/dev-tools/auth/oauth-m2m).

### Using `sling conns`

{% code overflow="wrap" %}

```bash
$ sling conns set DATABRICKS type=databricks host=<workspace-hostname> token=<access-token> warehouse_id=<warehouse-id> copy_method=zerobus client_id=<client-id> client_secret=<client-secret> zerobus_endpoint=https://<workspace-id>.zerobus.<region>.cloud.databricks.com
```

{% endcode %}

### Sling Env File YAML

```yaml
connections:
  DATABRICKS:
    type: databricks
    host: <workspace-hostname>
    token: <access-token>
    warehouse_id: <warehouse-id>
    catalog: <catalog>
    schema: <schema>
    copy_method: zerobus
    client_id: <client-id>
    client_secret: <client-secret>
    zerobus_endpoint: https://<workspace-id>.zerobus.<region>.cloud.databricks.com
```

{% hint style="info" %}
Zerobus writes Delta tables only. If `table_format` is `iceberg`, Sling uses `delta` for this connection.
{% endhint %}

{% hint style="warning" %}
`ARRAY`, `MAP`, `STRUCT`, and `VARIANT` columns are not supported. Incoming columns must match the target table: extra source columns, or missing non-null target columns, will fail.
{% endhint %}

If you are facing issues connecting, give [`sling assist`](/sling-cli/ai/assist.md) a try, or reach out to us at <support@slingdata.io>, on [discord](https://discord.gg/q5xtaSNDvp) or open a Github Issue [here](https://github.com/slingdata-io/sling-cli/issues).


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.slingdata.io/connections/database-connections/databricks.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
