Connecting S3 on Clarisights
Connecting an Amazon S3 bucket to Clarisights lets you bring your own custom data — CSV, compressed CSV, or Parquet files — into the platform alongside your other channels. Once the connection is in place, we read files from a path you specify and load them on the schedule your CSM configures.
At a glance
| Connector type | Object store, read-only |
| Authentication | IAM role with cross-account trust to Clarisights' AWS account |
| Permissions needed | s3:GetObject, s3:ListBucket |
| File formats | CSV, CSV.GZ, Parquet |
| Refresh cadence | Set per pipeline (configured by your CSM) |
| Limited rollout | No |
Authentication options
S3 connections to Clarisights use an IAM role in your AWS account that trusts Clarisights' AWS account. We assume that role from our side to read objects in your bucket — your data never leaves your account except via the explicit S3 reads we perform during a pipeline run.
Using an IAM role (rather than handing over long-lived access keys) means you keep full control: you can revoke access at any time by editing the trust policy or detaching the policy from the role.
Setting up the connection
Step 1 — Create an IAM role in your AWS account
In the AWS Console, go to IAM → Roles → Create role. Choose Custom trust policy and add a trust policy that allows Clarisights' AWS account to assume this role. Your CSM will provide our exact AWS account ID (and an External ID, if your security policy requires one) — paste it into the Principal.AWS field of the trust policy.
[TODO: Screenshot needed — AWS IAM Create role screen with Custom trust policy selected]
⚠ Common error:
AccessDenied: User is not authorized to perform sts:AssumeRole— usually caused by a typo in the AWS account ID or a missing External ID in the trust policy. Double-check the values your CSM provided.
Step 2 — Attach an S3 access policy to the role
Attach an inline IAM policy (or a managed policy) that grants s3:GetObject on the objects we need to read and s3:ListBucket on the bucket itself. Scope the policy to the specific bucket — and ideally the specific prefix — that holds the data you want Clarisights to read.
Example minimal policy:
{ "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": ["s3:ListBucket"], "Resource": "arn:aws:s3:::YOUR-BUCKET-NAME" }, { "Effect": "Allow", "Action": ["s3:GetObject"], "Resource": "arn:aws:s3:::YOUR-BUCKET-NAME/YOUR-PREFIX/*" } ] }
⚠ Common error:
AccessDeniedwhen we try to read a file — usually means the role's policy is missings3:GetObject, the resource ARN is wrong, or your bucket policy explicitly denies the role. Verify the bucket name and prefix in the policy match exactly.
Step 3 — Share the connection details with your CSM
Once the role is in place, send your CSM the following:
Bucket name
Bucket region
IAM role ARN (from the role's Summary page in IAM)
The path or prefix where your data files will land (for example
exports/daily/, or a date-templated path likeexports/<year>/<month>/<day_of_month>/— see Path date placeholders)The file format (CSV, CSV.GZ, or Parquet)
Step 4 — Drop your files into the bucket
Place the files at the agreed path. Once the pipeline is configured on our side, Clarisights will start reading from that path on the cadence your CSM set up.
⚠ Common error:
NoSuchKey— Clarisights tried to read a file at the configured path but didn't find it. Check that the prefix is correct and that your producer is actually writing files there.
Connection details exchange
| You provide to Clarisights | Clarisights provides to you |
| Bucket name | Clarisights AWS account ID (for the role's trust policy) |
| Bucket region | External ID, if required by your security policy |
| IAM role ARN | |
File path or prefix (may include <year>, <month>, <day_of_month> placeholders) | |
| File format (CSV, CSV.GZ, or Parquet) |
What we read
CSV (
.csv) — the first row is treated as the header. Subsequent rows are read as data.Compressed CSV (
.csv.gz,.gz) — same as CSV, decompressed on read.Parquet (
.parquet) — column names and types are taken from the Parquet schema. If your data is partitioned by date (for exampleorder_date=2026-04-01/), Clarisights can read only the partitions needed for the current pull window, which is much faster on large historical buckets.
If multiple files exist under the configured prefix, we read all of them and concatenate the rows. The schema must be consistent across files — see the limitations below.
For CSV-based pipelines, either keep the filename static and overwrite it on each export, or include a predictable date pattern in the filename (for example orders_2026-04-25.csv) so the pipeline can pick up the right file each run.
Path date placeholders
If your data is written to S3 in a time-partitioned layout (a folder per day, per month, etc.), you can include date placeholders in the path or prefix you give your CSM. Clarisights resolves these placeholders on every pipeline run, using the channel's configured timezone, so each run reads the folder for the date being pulled.
Supported tokens:
<year>— 4-digit year (for example2026)<month>— 2-digit month (for example04)<day_of_month>— 2-digit day of month (for example25)*::<today-Ndays>— append this to the path to resolve the date placeholders for N days ago instead of today's date. Useful when your export for a given day lands the next morning, or when you want a small look-back window.
Example. A daily export lands at s3://my-bucket/exports/2026/04/25/orders.csv. Configure the path as:
exports/<year>/<month>/<day_of_month>/orders.csv
On 25 Apr 2026, Clarisights resolves this to exports/2026/04/25/orders.csv. On 26 Apr 2026, it resolves to exports/2026/04/26/orders.csv, and so on. If your producer writes "yesterday's" file the next morning, append *::<today-1days> to read the previous day's folder on each run.
Date placeholders are always resolved in the timezone configured for the channel — so if your channel timezone is set to Europe/Berlin, midnight Berlin time is when the day rolls over.
Connector specifics
Glacier-tier objects must be restored to a standard storage class before Clarisights can read them. We don't trigger restores automatically — restore them in S3 first, then the next pipeline run will pick them up.
Cross-region transfers incur AWS data transfer fees on your side. Clarisights' AWS account may be in a different region from your bucket; the connection still works, but expect the usual cross-region egress charges.
Bucket policies on your side must not
Denythe IAM role we assume. An explicit deny in a bucket policy will override the IAM role's allow and will surface asAccessDeniedon our reads.
Limitations & known constraints
The schema must be consistent across files in a single prefix. If column names or types drift between files, the pipeline will fail or load nulls. Keep a stable schema, or split distinct schemas into distinct prefixes / pipelines.
Very large individual files extend the pull's runtime. If you control the export, prefer many small-to-medium files over one very large file — this also makes Parquet partitioning more useful.
Parquet without date partitioning still works, but Clarisights has to scan the full prefix on every run. For multi-year datasets, partitioning by day on your primary date column is strongly recommended.
The connector is read-only — Clarisights does not write back to your bucket.
Operating notes
Refresh cadence is configured per pipeline by your CSM. Common cadences are hourly and daily; pick what matches your file-drop schedule.
Failed pulls (for example, missing files or
AccessDenied) surface as integration errors in Clarisights and can be reviewed with your CSM or via support.You can rotate or restrict access at any time by editing the role's trust policy or detaching the S3 policy. Let your CSM know if you make changes that affect the role ARN or trust relationship so we can update our side in step.
If you change the bucket name, region, or file path, share the new details with your CSM — these are stored on our pipeline configuration and need to be updated explicitly.
Need help?
When contacting support from the in-app messenger, please include:
The integration name and account ID (Integrations → Channel)
The exact error message or screenshot
The step where the issue occurred
When the issue started