Routy

Data Export

Pull your raw clicks, page views and impressions out as CSV, JSON or Parquet, and any report you've built as CSV or Excel. Event exports run as a job you can queue, watch and cancel.

What this feature does

There are two kinds of export in Routy, and they answer different questions.

An event export gives you the rows behind the numbers: one record per click, per page view or per impression, with the IP, user agent, referrer, geo, sub-ids, tracker and destination that each one carried. You ask for a source, a date range and a format, and Routy runs it as a background job. That's what you want for a notebook, a warehouse load or a partner who needs to audit what you billed them for.

A report export gives you a report you already built, as a file. You pick the report in Report builder, set your columns, filters and sort, and export it to CSV or Excel. That's what you want when the answer is a table rather than a row dump.

The two share their vocabulary, so what you learn on one carries over: the same filters, the same offset and limit paging, a pre-signed link at the end. Where they differ is in the waiting. An event export is a job you queue, watch and cancel. A report export answers in the response, with the link in it.

What you'll get out of it

  • Clicks, page views and impressions at the event level. A click export carries the click id, sequence, affiliate, traffic source, server and client time, client IP, referrer, request headers, user agent, request and target URL, sub-id A and B, country, region and city, the brand link and account, the external click id, the affiliate tracker and the outgoing dynamic parameter. It's the raw row, not a summary of it.
  • CSV, JSON or Parquet on an event export, with CSV as the default. CSV opens in Excel or Sheets and is what most people want. JSON suits a script that reads the file directly. Parquet is typed and columnar, which is what you want when the file is going into pandas, DuckDB or a lake.
  • CSV or Excel on a report export, defaulting to CSV. The export is capped at 100,000 rows and runs with column totals switched off, so a report with a large result set comes back truncated rather than slow. The account-integration event exports go further and tell you how many rows were available next to how many were written, so a truncated file is visible rather than something you discover later.
  • A job you can follow. An event export moves through Queued, InProgress, Completed, Failed or Cancelled, and the record keeps createdAt, startedOn, endedOn and the number of records exported. Cancelling a queued job is immediate. Cancelling a running one takes effect at the next day boundary, since the export checks its own status between days.
  • Day-by-day processing with historical days cached. Every day except today is kept in object storage per account, source and format, so a second export over an overlapping range reuses what's already prepared. Today is never cached, because it's still filling up.
  • A manifest on every event export recording what's in it: the job id, the account, the date range, the format, when it ran, the total record count, and for each day whether that day came from cache or was generated fresh. Two exports of the same source and range are told apart by their manifests rather than by guessing from the file names.
  • An optional callbackUrl on the request. Routy POSTs to it when the job finishes, so a pipeline can wait on the export instead of polling for it.
  • Uploads that survive a wobble. Each batch of rows is retried up to ten times with exponential backoff before the day is treated as failed.

Worth knowing before you plan around it. An event export can't start more than 120 days ago, so Routy's export endpoint isn't the way to backfill a warehouse from the beginning of time; that's what the warehouse connector is for. Conversions aren't an event export source today, so conversion data comes out through the report export or the performance report endpoint. filterCriteria only narrows a page-views export, and only by TrafficSourceId; a clicks or impressions export gives you the whole day for the account. And there's no scheduled or recurring export: each one is a request you make.

How it actually works

Running an event export

POST /v1/exports with the source, the range and the format:

{
  "exportSource": "PageViews",
  "from": "2026-09-01T00:00:00Z",
  "to": "2026-09-28T23:59:59Z",
  "format": "CSV",
  "filterCriteria": { "TrafficSourceId": 5 }
}

from later than to, a from more than 120 days old, a to in the future, or a format other than CSV, JSON or Parquet all come back as 400 with the reason. Asking for an export that matches one already running or already finished returns 409 carrying that existing job, so a double-click on the button sends you to the file you already have instead of running the same range a second time.

From there, GET /v1/exports lists your jobs and filters by ExportSource and Status with Offset and Limit (default 20), GET /v1/exports/{id} reads one, and POST /v1/exports/{id}/cancel stops a job that is queued or running. You only see your own account's jobs. Admin-created exports run at high priority; an export you create is queued at normal priority, so it may wait a little before it starts.

Day caching, and why your second export is quick

Routy walks the range one day at a time. For every day except today, it looks for a prepared copy for your account, that source and that format. If one exists, it's reused. If not, the day is read from the event store in batches of 10,000 rows, each batch written as its own numbered file, and the result is kept for next time.

So the cost of an export falls as you iterate. The first pass over 30 days of clicks pays for 30 days; change a column in your analysis and re-export the same range and you pay for today only. Most exports finish between 30 seconds and 5 minutes. Ninety days of clicks on a high-volume account can take up to half an hour on the first run.

Downloading

GET /v1/exports/{id}/download returns a record count and a map of file names to pre-signed URLs. Because days and batches are written separately, an export is a set of files rather than one, named by their day and whether that day was cached or fresh, like 2026-09-14_cached_001.csv. The URLs are short-lived by design. When one has expired, call the download endpoint again for fresh ones; nothing has to be re-exported.

Running a report export

POST /v1/report-builder/{report}/export takes the same body as the report's data request plus format, which is csv or excel. Paging on the request is replaced with the export cap and column totals are turned off, then the file is generated and uploaded. The response carries a pre-signed downloadUrl good for a day, a fileName of the report name plus a timestamp, the row count and the format.

The account-integration event exports work the same way, at POST /v1/integrations/accounts/events/days/export and its siblings. dataSource decides what the rows are: Raw writes one row per event, Pivot writes the per-day pivoted columns. Those responses carry the number of rows available as well as the number written, so you can see when the 100,000-row cap bit.

Why this is worth doing

The export is what you reach for when the question doesn't fit a report. A partner wants a CSV they can reconcile against their own numbers. An analyst wants the click rows in a notebook, joined against spend data that lives in the ad platform. Somebody asks why Tuesday looks wrong and the answer is in the user agents.

Day caching is the part that changes how people work. The obvious design is to run the whole range on every request, which means every change to an analysis costs a full re-export. Here only today is ever recomputed, so the second and third pass over the same window finish in a fraction of the time. Analysts who spend real time on Routy data notice this in the first week and stop batching their questions.

The format choice is worth a moment too. CSV is right for a partner handoff and for anything a person will open. Parquet is right for a notebook, because the columns are typed and compressed and load without a schema guess. JSON is the one to pick when a script reads the file and you'd rather not deal with CSV quoting. If you're unsure, take CSV.

What the export isn't is an archive. The 120-day floor on from means the export endpoint covers recent operations, not history, and a warehouse that needs everything should be fed by the connector instead.

Frequently asked questions

What can I export at the event level?

Clicks, page views and impressions. Conversions aren't an event export source today; conversion data comes out of a report export or the performance report endpoint instead.

How far back can I go?

120 days. A from date older than that is refused with "'From' date cannot be older than 120 days". A to date in the future is refused too.

Is there a row or size limit on an event export?

No cap on the number of event rows. The range is the limit, and the file is split: 10,000 rows per file, numbered within each day. A report export is capped at 100,000 rows.

Which format should I pick?

CSV unless you have a reason. Parquet when the file is going into a notebook, a warehouse or a lake. JSON when a script reads it and you want the nesting kept.

How do I know when it's finished?

Poll GET /v1/exports/{id} for the status, or pass a callbackUrl when you create the job and Routy POSTs to it on completion. There's no email notification.

Can I cancel an export that's already running?

Yes. POST /v1/exports/{id}/cancel. A queued job stops at once. A running job stops at the next day boundary, because it rereads its own status between days, so a long export can take a moment to come to a halt.

How long does the download link last?

The pre-signed links on an event export are short-lived. If one stops working, call the download endpoint again and you get fresh links for the same files. A report export's link is good for a day.

Can I schedule a recurring export?

Not yet. Every export is a request you make, through the UI or the API. If you need data arriving somewhere on a schedule, that's the warehouse connector rather than this.

Why was my second export of the same range so much faster?

Every historical day was already prepared and was reused. Only today is ever recomputed. The manifest on each export says, day by day, which days came from cache and which were generated fresh.

Can I export only one traffic source?

On a page-views export, yes, with filterCriteria: { "TrafficSourceId": <id> }. Clicks and impressions exports return the whole day for your account, so filter those after download.

Ready to try Data Export?

Open Exports, pick the source, the date range and the format, and submit. The job appears in the list with its status, and the download becomes available when it finishes.