> For the complete documentation index, see [llms.txt](https://docs.slingdata.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.slingdata.io/connections/datalake-connections/lancedb.md).

# LanceDB

Connect & Ingest data from / to a LanceDB namespace

LanceDB is a lakehouse and vector database built on the [Lance](https://lance.org/) columnar format. Sling connects to a LanceDB **namespace**: a directory (local or object store) that holds one `<table>.lance` dataset per table. See <https://lancedb.com/> for more details.

The namespace is served through DuckDB's [`lance` extension](https://duckdb.org/docs/current/core_extensions/lance.html), which Sling attaches as a DuckDB catalog. The DuckDB SQL dialect, type mapping and merge engine therefore apply, and Sling needs no extra dependency of its own: DuckDB installs and loads the `lance` extension on first use.

## Setup

The following credentials keys are accepted:

* `path` **(required)** -> The namespace root: the directory holding the `.lance` datasets. A local path (`/data/lancedb`, `./lancedb`) or an object store URI (`s3://bucket/lancedb`, `gs://bucket/lancedb`, `az://container/lancedb`). `instance` is accepted as an alias.
* `schema` (optional) -> The default schema to read/write. Default is `main`.
* `duckdb_version` (optional) -> The CLI version of DuckDB to use. You can also specify the env. variable `DUCKDB_VERSION`.
* `copy_format` (optional *v1.6.4+*) -> the data format that Sling uses to move data between Sling and DuckDB, for reads and writes. Acceptable values: `arrow`, `csv`. Default is `arrow`. Arrow needs the DuckDB `arrow` community extension. If DuckDB cannot load this extension (for example, with no internet access), Sling uses `csv`. **If you have issues with arrow, set `csv`**.
* `copy_method` (deprecated) -> use `copy_format`. `arrow_http` becomes `arrow`. All other values become `csv`.
* `max_buffer_size` (optional) -> the max buffer size to use when piping data from the DuckDB CLI. Specify this if you have extremely large one-line text values in your dataset. Default is `10485760` (10MB).
* `max_line_size` (optional) -> the max line size (in bytes) to use when piping CSV data into DuckDB. Sling raises this to `268435456` (256MB) when the data has string, text, json or binary columns, and uses `2000000` (2MB) otherwise. Specify this only if a single row is larger than 256MB. Applies when `copy_format` is `csv`.

{% hint style="info" %}
`read_only` is **not supported**. The DuckDB CLI cannot launch its in-memory database read-only, and the lance extension rejects a `READ_ONLY` attach. Sling returns an error if the property is set.
{% endhint %}

### Storage Configuration

The lance extension resolves `s3://` (plus the `s3a://` / `s3n://` aliases), `gs://` and `az://` (plus the `abfss://` alias). For S3-compatible services (MinIO, Cloudflare R2, Backblaze B2, ...) use an `s3://` path together with `s3_endpoint`.

For **S3/S3-compatible storage**:

* `s3_access_key_id` (optional) -> AWS access key ID
* `s3_secret_access_key` (optional) -> AWS secret access key
* `s3_session_token` (optional) -> AWS session token
* `s3_region` (optional) -> AWS region
* `s3_endpoint` (optional) -> S3-compatible endpoint URL (e.g. `http://localhost:9000` for MinIO)

When no keys are provided, the upstream SDK credential chain is used (environment variables, shared config/credentials profiles, instance metadata).

For **Azure Blob / ADLS Storage**:

* `azure_account_name` (optional) -> Azure storage account name
* `azure_account_key` (optional) -> Azure storage account key
* `azure_sas_token` (optional) -> Azure SAS token (used when no account key is given)

For **Google Cloud Storage**:

* `gcs_key_file` (optional) -> Path to a GCS service account key file

### Using `sling conns`

Here are examples of setting a connection named `LANCEDB`. We must provide the `type=lancedb` property:

{% code overflow="wrap" %}

```bash
# Local namespace
$ sling conns set LANCEDB type=lancedb path=/data/lancedb

# S3 namespace
$ sling conns set LANCEDB type=lancedb path=s3://my-bucket/lancedb \
    s3_access_key_id=AKIAIOSFODNN7EXAMPLE \
    s3_secret_access_key=wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY \
    s3_region=us-east-1

# MinIO / S3-compatible (plain HTTP)
$ sling conns set LANCEDB type=lancedb path=s3://my-bucket/lancedb \
    s3_endpoint=http://localhost:9000 \
    s3_access_key_id=minioadmin \
    s3_secret_access_key=minioadmin \
    s3_region=us-east-1

# Azure namespace
$ sling conns set LANCEDB type=lancedb path=az://my-container/lancedb \
    azure_account_name=<azure_account_name> \
    azure_account_key=<azure_account_key>

# GCS namespace
$ sling conns set LANCEDB type=lancedb path=gs://my-bucket/lancedb \
    gcs_key_file=/path/to/service-account.json

$ sling conns test LANCEDB
```

{% endcode %}

### Environment Variable

See [here](https://docs.slingdata.io/connections/datalake-connections/pages/eAdVs2BHCgdr6RS8GoJC#dot-env-file-.env.sling) to learn more about the `.env.sling` file.

{% code overflow="wrap" %}

```bash
# local namespace
export LANCEDB='{ type: lancedb, path: "/data/lancedb" }'

# S3 namespace
export LANCEDB='{
  type: lancedb,
  path: "s3://my-bucket/lancedb",
  s3_access_key_id: "AKIAIOSFODNN7EXAMPLE",
  s3_secret_access_key: "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY",
  s3_region: "us-east-1"
}'
```

{% endcode %}

### Sling Env File YAML

See [here](https://docs.slingdata.io/connections/datalake-connections/pages/eAdVs2BHCgdr6RS8GoJC#sling-env-file-env.yaml) to learn more about the sling `env.yaml` file.

```yaml
connections:
  LANCEDB:
    type: lancedb
    path: /data/lancedb
    schema: main

    # S3 Configuration (if using S3 storage)
    s3_access_key_id: <s3_access_key_id>
    s3_secret_access_key: <s3_secret_access_key>
    s3_region: <s3_region>
    s3_endpoint: http://localhost:9000

    # Azure Configuration (if using Azure storage)
    azure_account_name: <azure_account_name>
    azure_account_key: <azure_account_key>

    # GCS Configuration (if using GCS storage)
    gcs_key_file: <gcs_key_file>
```

## Common Usage Examples

### Basic Operations

```bash
# Query data
sling run --src-conn LANCEDB --src-stream "SELECT * FROM main.orders LIMIT 10" --stdout

# Export to Parquet
sling run --src-conn LANCEDB --src-stream main.orders --tgt-object file://./orders.parquet
```

### Data Import/Export

```bash
# Import a CSV into a Lance dataset
sling run --src-stream file://./data.csv --tgt-conn LANCEDB --tgt-object main.new_data

# Import from PostgreSQL
sling run --src-conn POSTGRES_DB --src-stream public.customers --tgt-conn LANCEDB --tgt-object main.customers

# Replicate a Lance table into PostgreSQL
sling run --src-conn LANCEDB --src-stream main.orders --tgt-conn POSTGRES_DB --tgt-object public.orders
```

## Merge Strategies

All merge strategies are expressed as a single `MERGE INTO`, which the lance extension supports. The `delete ... where exists` and `update ... from` forms cannot be planned by the extension (`unsupported DELETE plan: expected 1 child, got 2`).

| Strategy              | Behavior                                                                                               |
| --------------------- | ------------------------------------------------------------------------------------------------------ |
| `insert`              | insert rows that are not matched by primary key                                                        |
| `update`              | update matched rows in place                                                                           |
| `update_insert`       | update matched rows, insert unmatched rows                                                             |
| `delete_insert`       | update matched rows in place (equivalent end state to delete + insert)                                 |
| `change_capture`      | apply `_sling_synced_op` (`I`, `U`, `D`) events, keeping the highest `_sling_cdc_seq` per primary key  |
| `change_capture_soft` | same, but marks deletes with `_sling_synced_at` / `_sling_synced_op = 'D'` instead of removing the row |

```yaml
source: POSTGRES_DB
target: LANCEDB

defaults:
  mode: incremental
  primary_key: id

streams:
  public.customers:
    object: main.customers
    update_key: updated_at
    target_options:
      merge_strategy: delete_insert
```

### Change Capture (CDC)

Replicating a source's change stream into LanceDB uses `change_capture`, with the events supplied in the `_sling_synced_op` and `_sling_cdc_seq` columns:

```yaml
source: POSTGRES_DB
target: LANCEDB

defaults:
  mode: incremental
  primary_key: id

streams:
  cdc_events:
    sql: |
      select 1 as id, 'Alice' as name, 'U' as _sling_synced_op, 4::bigint as _sling_cdc_seq
      union all select 2, 'Bob', 'D', 2
      union all select 3, 'Charlie', 'I', 3
    object: main.customers
    target_options:
      merge_strategy: change_capture
```

See [Capture Deletes](/examples/database-to-database/capture_deletes.md) for the `delete_missing` option, which is also supported for LanceDB targets.

## Limitations

* **Views are not persisted.** The lance extension keeps `CREATE VIEW` in the session that created it: the view is visible in `information_schema.tables` for that session, but nothing is written into the namespace and it is gone after reconnecting. The Lance namespace specification has no view operations either. Use a table (`CREATE TABLE ... AS SELECT`) when the result must survive a reconnection.
* **Object store schemes.** Only `s3://` (plus `s3a://` / `s3n://`), `gs://` and `az://` (plus `abfss://`) are resolved. Other schemes (e.g. `r2://`, `gcs://`, `abfs://`) are rejected by the extension; use the matching supported scheme with an explicit endpoint instead.
* **JSON columns are stored as text.** The extension writes `JSON` columns as `VARCHAR`, so a JSON column round-trips as text.
* **No secondary indexes.** Lance datasets have no secondary indexes, so Sling does not create indexes (including unique indexes) on LanceDB targets.

If you are facing issues connecting, give [`sling assist`](/sling-cli/ai/assist.md) a try, or reach out to us at <support@slingdata.io>, on [discord](https://discord.gg/q5xtaSNDvp) or open a Github Issue [here](https://github.com/slingdata-io/sling-cli/issues).


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.slingdata.io/connections/datalake-connections/lancedb.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
