Email Data Source
The Email connector allows you to extract structured data from email attachments. It monitors a dedicated inbox and automatically processes CSV files attached to incoming emails, converting them into a data stream that can be loaded into your destination.
Overview
This connector is ideal for scenarios where you receive regular data reports via email with CSV attachments, such as:
- Automated reports from partners or vendors
- Daily/weekly data exports from systems that only support email delivery
- Marketing performance reports sent via email
- Financial statements or transaction reports delivered as email attachments
Prerequisites
- A data provider or system that can send reports via email with CSV attachments (directly attached or inside a ZIP file)
Supported File Types
The Email connector processes the following attachment types:
| File Type | Description |
|---|---|
.csv | CSV files are parsed directly |
.zip | ZIP archives containing CSV files are automatically extracted and processed (including nested ZIPs) |
How It Works
- Scheduled Runs: On each connection run (based on your configured schedule), the connector checks the inbox for new emails
- Attachment Extraction: The connector identifies and extracts supported attachments from any emails received since the last run
- Data Parsing: CSV files are parsed, with each column becoming a field in the data stream. Column headers are automatically converted to snake_case (e.g.,
First Namebecomesfirst_name) - Schema Detection: The schema is automatically inferred from the CSV header row of the most recent email
- Incremental Processing: Only new emails (based on timestamp) are processed in subsequent runs
Source Setup Guide
-
In the Extract platform, navigate to Sources → Add Source → Email
-
Provide a Table Name for the resulting data stream (this will be slugified, e.g.,
My Tablebecomesmy-table) -
Click Save — a dedicated email address will be provisioned for your source
-
Copy the provisioned email address and configure your data provider to send reports to this address
- Use the exact provisioned address. The Email connector validates that the inbox address is uniquely registered to your source; if the address is reused or not uniquely associated, the source will fail to run.
-
Send a test email with a CSV attachment to verify the setup
Connection Setup Guide
Once you've connected the Email source to a destination, configure:
- Connection Pull Schedule: How frequently Extract checks for new emails and attachments.
- Destination-specific settings: Dataset name, target schema, etc. (depending on your destination).
- Schema Migration Policy: How Extract handles schema changes when new files introduce new/changed columns.
Notes:
- The inbox address used by this source must be uniquely registered to this Email source. If the same inbox is connected to multiple Email sources (or ownership is ambiguous), the connection will fail. If you run into this, re-save the source configuration or contact support.
Data Stream Fields
In addition to the columns from your CSV files, the connector automatically adds the following metadata fields:
| Field | Type | Description |
|---|---|---|
_email_sender | String | The email address of the sender |
_email_attachment_filename | String | The filename of the processed attachment (for files inside ZIPs, includes the full path, e.g., report.zip/data.csv) |
_email_timestamp | String (ISO 8601) | The timestamp when the email was received, in RFC3339 format (e.g., 2024-01-15T10:30:00Z) |
Email Data Source FAQ
Q: How is the data schema determined?
A: The connector uses the most recent email's CSV attachment to determine the schema. The first row of the CSV file is treated as column headers. Headers are automatically converted to snake_case (e.g., First Name becomes first_name, userID becomes user_id).
Q: What happens if CSV schemas change between emails?
A: The schema is expected to remain consistent. If a new email contains a CSV with different columns, the connection may fail based on your Schema Migration Policy settings.
Q: Are all data types treated as strings?
A: Yes, all values from CSV files are imported as strings. You can apply transformations in your destination to cast values to appropriate data types.
Q: What happens with nested ZIP files?
A: The connector recursively processes ZIP files. If a ZIP contains another ZIP with CSV files inside, all CSVs will be extracted and processed. The _email_attachment_filename field will show the full path (e.g., outer.zip/inner.zip/data.csv).
Q: Are there file size limits?
A: Yes, individual attachments are limited to 100 MB. Attachments exceeding this limit will be skipped with an error logged.
Q: What email formats are supported?
A: The connector supports standard MIME-encoded emails with base64-encoded attachments. Most email providers and systems use this format by default.
Q: How do I send test emails?
A: Send an email with a CSV attachment to the email address provisioned for your source. The connector will process it on the next scheduled run. You can verify the setup by checking the connection run logs.
Q: What happens to files that aren't CSV or ZIP?
A: Attachments with unsupported file extensions (e.g., .xlsx, .pdf, .txt) are skipped. Only .csv and .zip files are processed.
Best Practices
- Consistent Schema: Ensure that all CSV files sent to the inbox have the same column structure
- Clear Naming: Use descriptive filenames for attachments to make it easier to identify data in the
_email_attachment_filenamefield - Keep Attachments Under 100 MB: Split large files if necessary to stay under the attachment size limit
- Use Standard CSV Format: Ensure the first row contains headers and data uses consistent delimiters
- Regular Monitoring: Check connection runs periodically to ensure emails are being processed correctly
Troubleshooting
| Issue | Possible Cause | Solution |
|---|---|---|
| Connection fails | Inbox address is not uniquely registered to this Email source | Re-save the source to ensure the inbox address is registered to the correct source. If the same inbox address is configured on multiple Email sources (or ownership is ambiguous), the connector will fail closed—use a single source per inbox address or contact support. |
| No data extracted | Attachment is not CSV or ZIP | Ensure attachments are in supported formats (.csv or .zip only). |
| No data extracted | No new emails since the last successful run (cursor) | Send a new email with a valid attachment, or verify the source’s incremental cursor/state if you expect newer messages to be processed. |
| Missing columns | Schema mismatch | Verify the CSV header row matches expected columns. Remember headers are converted to snake_case. |
| Column names don't match | Header case conversion | Column headers are converted to snake_case (e.g., firstName → first_name). |
| Duplicate data | Emails reprocessed | Ensure each email/object has a unique timestamp and verify the cursor/state if you see reprocessing. |
| Attachment skipped | File too large | Reduce attachment size to under 100 MB. |