Skip to navigation

Supports:

  • ✅ Models
  • ✅ Model sync destination
  • ✅ Bulk sync source
  • ✅ Bulk sync destination

Connection

Configuration

NameTypeDescriptionRequired
account_idstringAccount IDrequired
aws_access_key_idstringAccess Key ID with read/write access to a bucket.required
aws_secret_access_keystringSecret access keyrequired
bucket_namestringBucket name (folder optional); ex: polytomic/datasetrequired
csv_has_headersbooleanCSV files have headers

Whether CSV files have a header row with field names.
optional
is_single_tablebooleanFiles are time-based snapshots

Treat the files as a single table. ↓
optional

is_single_table = true

NameTypeDescriptionRequired
is_directory_snapshotbooleanMulti-directory multi-table ↓optional
{
"name": "Cloudflare R2 connection",
"type": "cloudflare_r2",
"configuration": {
"account_id": "",
"aws_access_key_id": "AKIAIOSFODNN7EXAMPLE",
"aws_secret_access_key": "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY",
"bucket_name": "polytomic/dataset",
"csv_has_headers": true,
"is_directory_snapshot": false,
"is_single_table": true,
"single_table_file_format": "csv",
"single_table_name": "collection",
"skip_lines": 0
}
}
is_directory_snapshot

When is_directory_snapshot is true:

NameTypeDescriptionRequired
directory_glob_patternstringTables glob pathrequired
{
"name": "Cloudflare R2 connection",
"type": "cloudflare_r2",
"configuration": {
"account_id": "",
"aws_access_key_id": "AKIAIOSFODNN7EXAMPLE",
"aws_secret_access_key": "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY",
"bucket_name": "polytomic/dataset",
"csv_has_headers": true,
"directory_glob_pattern": "",
"is_directory_snapshot": true,
"is_single_table": true,
"single_table_file_formats": []
}
}

is_single_table = true and is_directory_snapshot = false

NameTypeDescriptionRequired
single_table_file_formatstringFile format

Accepted values: csv ↓, json, parquet
optional
single_table_namestringCollection nameoptional
single_table_file_format

When single_table_file_format is csv or single_table_file_formats contains csv:

NameTypeDescriptionRequired
skip_linesintegerSkip first lines

Skip first N lines of each CSV file.
optional

is_single_table = true and is_directory_snapshot = true

NameTypeDescriptionRequired
single_table_file_formatsarrayFile formats that may be present across different tablesoptional

Example

{
"name": "Cloudflare R2 connection",
"type": "cloudflare_r2",
"configuration": {
"account_id": "",
"aws_access_key_id": "AKIAIOSFODNN7EXAMPLE",
"aws_secret_access_key": "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY",
"bucket_name": "polytomic/dataset",
"csv_has_headers": true,
"is_single_table": false
}
}

Model Sync

Source

Configuration

NameTypeDescriptionRequired
file_formatstringFile format

Accepted values: csv, json, parquet
optional
keystringObject key

The key of the object in the bucket to read from.
optional
model_fromstringFiles

The model is generated from a single file or a multi-file archive. Accepted values: single_file, multi_file_archive
required
skip_linesintegerSkip first lines

Skip first N lines of each CSV file.
optional
subfolderstringSubfolder to read files from from (optional)optional

Example

{
...
"configuration": {
"file_format": "csv",
"key": "",
"model_from": "single_file",
"skip_lines": 0,
"subfolder": ""
}
}

Target

Cloudflare R2 connections may be used as the destination in a model sync.

All targets

Configuration
NameTypeDescriptionRequired
formatstringOutput format

Output file encoding. Accepted values: csv, json-doc, json, parquet
optional
Example
{
...
"target": {
"configuration": {
"format": "csv"
}
}
}

Bulk Sync

Source

Cloudflare R2 connections may be used as a bulk sync source. No additional configuration options are required.

Destination

Configuration

NameTypeDescriptionRequired
advancedobjectoptional
formatstringOutput file encodingoptional
subfolderstringSubfolder to write to (optional)optional

Example

{
...
"destination_configuration": {
"advanced": {
"write_incremental": false
},
"format": "csv",
"subfolder": "reports"
}
}