Skip to content
Clarisights Knowledge Center home
InboxAsk a human

Connecting Google Cloud Storage on Clarisights

Google Cloud Storage (GCS) is a read-only object-store connector. Use it to land CSV, CSV.GZ, or Parquet files in a bucket and have Clarisights pick them up as a custom data source for reporting.

At a glance

Connector typeObject store (read-only)
AuthenticationService account / IAM grant
Permissions neededStorage Object Viewer on the bucket
File formatsCSV, CSV.GZ, Parquet
Refresh cadenceSet per pipeline
Limited rolloutNo

Authentication options

Clarisights reads from your bucket using a managed Google service account. There are no API keys or OAuth flows to manage on your side — you grant our service account IAM access to the bucket (or a path prefix within it), and we read from there.

  • IAM grant to a Clarisights-managed service account — grant Storage Object Viewer at the bucket or prefix level. We will share the service account email during onboarding.

Setting up the connection

Step 1 — Identify or create the bucket

In the Google Cloud Console, choose an existing bucket or create a new one to hold the files you want Clarisights to read. Decide on a path or prefix that will hold the dataset (for example gs://my-bucket/clarisights/orders/).

Step 2 — Grant access to the Clarisights service account

Your Clarisights contact will share a service account email (it ends in .gserviceaccount.com). In the Cloud Console, open the bucket, go to Permissions, and grant that service account the Storage Object Viewer role. You can scope the grant to the whole bucket or to a prefix using IAM Conditions.

[TODO: Screenshot needed — GCS bucket Permissions tab with Storage Object Viewer grant]

⚠ Common error: "Access denied" or HTTP 403 when Clarisights tries to read the bucket → the service account is missing the Storage Object Viewer role on the bucket or prefix → re-check the IAM binding on the bucket.

Step 3 — Share the connection details with Clarisights

Send your Clarisights contact the bucket name, the path or prefix, and the file format (CSV, CSV.GZ, or Parquet). We will configure the pipeline and confirm when the first read succeeds.

⚠ Common error: "Object not found" or empty reads → the path or prefix is wrong, or files are nested under a deeper folder than configured → verify the exact gs://<bucket>/<prefix>/ URI in the Cloud Console and resend it.

Connection details exchange

You provideClarisights provides
Bucket name (e.g. my-bucket)Service account email to grant access to
File path or prefix (e.g. clarisights/orders/) (may include <year>, <month>, <day_of_month> placeholders)Confirmation once the first read succeeds
File format (CSV, CSV.GZ, or Parquet)Pipeline configuration and refresh cadence
IAM grant of Storage Object Viewer on the bucket or prefix—

What we read

  • Supported formats: CSV, gzip-compressed CSV (.csv.gz), and Parquet.

  • CSV header convention: the first row of every CSV file is treated as the header. Column names in that row become the field names in Clarisights.

  • Multiple files in a prefix: every file under the configured path or prefix is read as part of one logical dataset and concatenated. New files dropped into the prefix are picked up on the next refresh.

Path date placeholders

If your data is laid out by date (a common pattern for daily exports), the file path or prefix can include date placeholders that Clarisights resolves at pull time using the channel's configured timezone. This means you configure the path once, and each scheduled run reads the folder for the appropriate date automatically.

Supported tokens:

  • <year> — 4-digit year (e.g. 2026)

  • <month> — 2-digit month (e.g. 04)

  • <day_of_month> — 2-digit day (e.g. 25)

  • *::<today-Ndays> — resolves the date placeholders for N days ago instead of today (useful when files are published with a delay).

Example: if your daily exports land at gs://my-bucket/exports/2026/04/25/orders.csv, configure the path as exports/<year>/<month>/<day_of_month>/orders.csv. On each run, Clarisights substitutes the current date (in the channel's timezone) and reads that day's file. To always read the previous day's file, append *::<today-1days> so the placeholders resolve for yesterday.

Connector specifics

  • Storage class matters: Standard and Nearline buckets read with no extra latency. Coldline and Archive objects have a retrieval lag and incur additional retrieval fees on the Google side — we recommend Standard or Nearline for actively-refreshed datasets.

  • Schema must match across files in a prefix: when multiple files sit under one prefix, they are read as a single dataset, so column names and types must be consistent across all files. Mixing schemas in the same prefix will fail the read.

  • Path conventions: files at gs://<bucket>/<prefix>/... are read as one logical dataset. Use separate prefixes for separate datasets rather than mixing them in one folder.

Limitations & known constraints

  • Read-only connector — Clarisights does not write back to your bucket.

  • Only CSV, CSV.GZ, and Parquet are supported. Other formats (JSON, Avro, ORC, Excel) are not read.

  • All files under a configured prefix must share the same schema. Schema drift across files in the same prefix will fail the read.

  • Coldline and Archive objects are readable but slower and cost extra to retrieve; not recommended for active pipelines.

  • Bucket-level encryption with customer-managed keys (CMEK) is supported as long as the Clarisights service account has the necessary key access; customer-supplied encryption keys (CSEK) are not supported.

Operating notes

  • Refresh cadence is set per pipeline during onboarding. New files dropped into the prefix between refreshes are picked up on the next scheduled run.

  • IAM rotation: if you ever need to revoke access, remove the Storage Object Viewer binding for the Clarisights service account on the bucket; reinstating it restores reads on the next refresh.

  • Adding new files or prefixes: drop new files under the existing prefix to extend the dataset, or contact Clarisights to onboard an additional prefix as a separate source.

Need help?

When contacting support from the in-app messenger, please include:

  • The integration name and bucket name

  • The exact error message or screenshot

  • The step where the issue occurred

  • When the issue started