> For the complete documentation index, see [llms.txt](https://docs.slingdata.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.slingdata.io/connections/database-connections/opensearch.md).

# OpenSearch

Extract data from an OpenSearch instance

[OpenSearch](https://opensearch.org/) is the community-driven, Apache-2.0 licensed search and analytics suite (a fork of Elasticsearch). Sling connects to OpenSearch using the official [`opensearch-go`](https://github.com/opensearch-project/opensearch-go) client and extracts documents from one or more indexes.

{% hint style="info" %}
Sling is **read-only** for OpenSearch: you can extract data *from* an OpenSearch instance, but writing to an OpenSearch index as a target is not supported.
{% endhint %}

## Setup

The following credentials keys are accepted:

* `http_url` (optional) -> Comma-separated list of HTTP URLs to connect to the cluster
* `host` (optional) -> The hostname / ip of the instance
* `port` (optional) -> The port of the instance. Default is `9200`
* `user` (optional) -> The username to access the instance
* `password` (optional) -> The password to access the instance
* `tls` (optional) -> whether to use TLS for connecting (`true`/`skip-verify`/`custom`)
* `cert_file` (optional) -> the client certificate for TLS (file path or raw)
* `cert_key_file` (optional) -> the client key for TLS (file path or raw)
* `cert_ca_file` (optional) -> the client CA certificate for TLS (file path or raw)

### Using `sling conns`

Here are examples of setting a connection named `OPENSEARCH`. We must provide the `type=opensearch` property:

{% code overflow="wrap" %}

```bash
$ sling conns set OPENSEARCH type=opensearch host=<host> user=<user> password=<password> port=9200
```

{% endcode %}

### Environment Variable

See [here](https://docs.slingdata.io/connections/database-connections/pages/eAdVs2BHCgdr6RS8GoJC#dot-env-file-.env.sling) to learn more about the `.env.sling` file.

{% code overflow="wrap" %}

```bash
export OPENSEARCH='opensearch://user:password@host.ip:9200'
export OPENSEARCH='{ type: opensearch, user: "user", password: "password", host: "host.ip", port: 9200 }'
```

{% endcode %}

### Sling Env File YAML

See [here](https://docs.slingdata.io/connections/database-connections/pages/eAdVs2BHCgdr6RS8GoJC#sling-env-file-env.yaml) to learn more about the sling `env.yaml` file.

```yaml
connections:
  OPENSEARCH:
    type: opensearch
    host: host.ip
    port: 9200
    user: <user>
    password: <password>
    tls: true # if TLS connection needed

  OPENSEARCH_URL:
    type: opensearch
    http_url: http://host1:9200,http://host2:9200
```

## Indexes and Tables

Each OpenSearch index is exposed as a stream. The index name is used directly as the stream name:

```yaml
streams:
  my_index:
  logs-2024.01.01:
```

Columns are derived from the index mapping: `text`, `keyword` and `string` fields become `text` columns, `long` and `integer` fields become `bigint`, `float` and `double` become `float`, `date` becomes `timestamp`, `boolean` becomes `bool`, and `binary` becomes `binary`. When discovering columns from the index mapping, nested `object` fields are expanded into dotted column names (for example `meta.city`). If an index has no mapping, a single `data` (`json`) column is returned.

## Replication

Here's an example replication configuration to sync OpenSearch indexes to a PostgreSQL database:

```yaml
source: OPENSEARCH
target: MY_POSTGRES

defaults:
  mode: full-refresh
  object: public.{stream_table}

streams:
  # extract an entire index
  my_index:

  # extract a limited number of documents
  events:
    source_options:
      limit: 10000

  # incremental extraction based on a date or numeric field
  logs:
    mode: incremental
    primary_key: id
    update_key: created_at
```

### Flattening Documents

By default, each document is extracted as a single `data` JSON column:

```csv
data
"{""amount"":10.5,""created"":""2024-01-01T00:00:00Z"",""id"":1,""name"":""alice""}"
```

Set the `flatten` source option to expand the document fields into individual columns. Nested objects are flattened using `__` as the separator, and arrays are kept as JSON:

```yaml
streams:
  my_index:
    source_options:
      flatten: true
```

```
amount,created,id,name,meta__city,meta__zip,tags
5.5,,10,dave,paris,75001,"[""a"",""b""]"
```

### Incremental and Backfill

For incremental extraction, set `update_key` to a date or numeric field in the index. Sling reads documents where that field is greater than the last stored value:

```yaml
streams:
  logs:
    mode: incremental
    primary_key: id
    update_key: created_at
```

To backfill a specific range of an index, use the `range` source option together with `update_key`:

```yaml
streams:
  logs:
    source_options:
      range: "2024-01-01,2024-06-01"
    primary_key: id
    update_key: created_at
```

If you are facing issues connecting, give [`sling assist`](/sling-cli/ai/assist.md) a try, or reach out to us at <support@slingdata.io>, on [discord](https://discord.gg/q5xtaSNDvp) or open a Github Issue [here](https://github.com/slingdata-io/sling-cli/issues).


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.slingdata.io/connections/database-connections/opensearch.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
