# Collect supplier CSV exports with n8n

Download **Supplier-CSV-Collector.n8n.json**, or **Supplier-n8n-Collector.zip** for the workflow, this guide and its offline tests, from the [supplier guide](https://supplier-pdf-to-csv.maag8484.chatgpt.site/#automation).

Use this collector after starting a Supplier Price List to CSV run through Apify Console or the existing Python client. It waits for that exact run and produces two downloadable files: `PRODUCTS.csv` and `DIAGNOSTICS.json`. The diagnostics keep the summary, review records and document outcomes together.

**Validation status:** 32 offline checks passed on September 14, 2026. The tests execute the workflow's embedded JavaScript and expressions against simulated node responses. The JSON has not been imported or executed in a customer n8n instance. Start with an existing fictional demo run before connecting a destination.

## Import and collect

1. In your n8n workflow editor, use **Import from File** and select `Supplier-CSV-Collector.n8n.json`.
2. Open **Run ID** and replace `PASTE_EXISTING_RUN_ID` with the 17-character ID of your existing extraction run. Find it in Apify Console or the Python client's saved state. Do not use the Actor ID, task ID or a Console URL. This template supports the extraction Actor only.
3. Create a protected **Header Auth** credential: name `Authorization`, value `Bearer YOUR_APIFY_TOKEN`. Select it on all five HTTP Request nodes: Read existing run, Read SUMMARY, Read REVIEW, Read DOCUMENTS and Read PRODUCTS.csv. Keep the token out of code, query strings, notes and exported workflow files. Use the narrowest credential permissions that allow reading your runs and their output records.
4. Execute the workflow manually. It checks run status at most 18 times, with 10 seconds between pending checks. Request failures stop collection. If the workflow ends while the run is still pending, execute collection again with the same ID later.
5. Open **Validated export**. Download the `products` binary as PRODUCTS.csv and `diagnostics` as DIAGNOSTICS.json. CSV product count must match SUMMARY's delivered row count before either export is returned.

Importing does not activate a schedule, start an Actor or create a subscription. Executing this collector performs authenticated API reads in your own Apify account; API/storage usage and n8n execution limits may apply. It cannot repeat the extraction's product-row charge because it contains no Actor-start request. It does not stop a cloud run when collection times out.

## Use the output safely

CSV count agreement establishes delivery consistency, not completeness or extraction accuracy. The diagnostics preserve partial-output indicators, review records and document failures. `destinationWriteApproved` stays false: a destination-specific mapping and exception policy are still required before enabling downstream imports. Keep SKUs as text and hold unresolved or partial results automatically. No ERP/store write, inventory update or human correction service is included.

The CSV is preserved byte-for-byte. Formula-like spreadsheet text is flagged, not executed. Handle imported columns as data. A processing guard rejects CSV text above 25 MiB after download; it is not a network transfer cap. Protect n8n execution history and outputs if your supplier data is confidential. Configure retention and your own persistent destination; n8n execution history is not a permanent archive.

## Recover without another extraction

- HTTP errors, missing records, actor/run mismatches and malformed CSV stop the workflow. Keep the same run ID for recovery. Do not start a replacement run automatically.
- A failed, aborted or timed-out Actor run is reported as such; it is never restarted by this collector.
- The exported template has no credentials, pinned private data, webhook, schedule or downstream writes. Bind your own credential after importing.
- Collection may be repeated or executed concurrently without starting another extraction. If you later attach a destination, deduplicate exports there by **run ID + output key**, so retries do not duplicate destination records.
- To integrate into an existing automation, pass one saved run ID into Validate input in place of the manual Run ID node. Validate that integration in your own instance before activation. Keep the chargeable start step separate and protected by a durable document-revision reservation.

## Reproduce the offline checks

From the extracted collector ZIP, with Node.js 18 or newer:

```bash
node tools/test-n8n-collector.mjs
```

These tests make no network requests. They simulate n8n's scheduling, HTTP responses and binary-file helper; they do not test n8n's importer, your credentials, actual API availability or a destination integration.

## Starting a new document revision

The original Python kit remains available separately. It already reserves a persistent state file before a one-time start and can resume an existing run. Use it or Apify Console to start the initial compatible catalog. Keep an explicit positive run budget. Reprocessing a document in a separate run can be charged again.

The advanced setup recipe below explains how to implement the same separation when building your own start workflow. It is guidance, not an additional imported or activated workflow.

---

## Advanced start-once / collect-existing recipe

Use your own n8n instance and Apify account to start a catalog extraction, wait for the result and collect its CSV. This start-workflow recipe is separate from the downloadable collector; it has not been imported or executed in your n8n account. No workflow is activated by downloading this kit.

The [live Actor](https://apify.com/cory8484/my-actor) and the included sample let you inspect the output first. This workflow is useful when compatible supplier PDFs arrive on recurring catalog updates. It does not discover private supplier documents or perform ERP writes.

## Start with two separate workflows

Separating the one-time start from collection makes it possible to recover a slow run without repeating the chargeable request.

**A: Start a document revision.** Use Manual Trigger during setup. Place a durable supplier/revision reservation before the HTTP request. If your automation store cannot atomically reserve a unique revision, keep this manual until that safeguard exists. Save the returned run ID against the revision immediately. Do not loop back to the start node.

**B: Collect that existing run.** Supply the saved run ID. Check status, wait while pending, and download records only after success. Run B can resume independently of A.

Use n8n's protected Header Auth credential for the HTTP nodes: header name `Authorization`, value `Bearer YOUR_APIFY_TOKEN`. Configure the actual value privately. Do not place the token in node code, query parameters or shared workflow JSON. The [HTTP Request documentation](https://docs.n8n.io/integrations/builtin/core-nodes/n8n-nodes-base.httprequest/) covers authentication and response formats.

## HTTP request settings

| Step | Method and endpoint | Settings / handling |
|---|---|---|
| Start, once | POST `https://api.apify.com/v2/actors/HIlb5I6IftACciGRI/runs` | Query: `build=latest`, `maxTotalChargeUsd=0.25`, `timeout=120`, `memory=1024`, `restartOnError=false`. JSON body: `{}` for demo or your configured catalog input. Retry On Fail off; HTTP timeout 30000 ms. |
| Save identity | No API call | Persist `data.id` to the reserved supplier/revision. Confirm `data.actId` equals `HIlb5I6IftACciGRI` and `data.options.maxTotalChargeUsd` equals `0.25`. Stop on mismatch. |
| Check existing run | GET `https://api.apify.com/v2/actor-runs/RUN_ID` | Replace RUN_ID with that saved ID. Request JSON. Validate actor/run identity on each read. |
| Poll when pending | No start request | If status is READY, RUNNING, TIMING-OUT or ABORTING, use a 10-second Wait and return to the GET. Bound the loop to 18 checks. |
| Collect records | GET `https://api.apify.com/v2/key-value-stores/STORE_ID/records/KEY` | Take STORE_ID from the successful run's `data.defaultKeyValueStoreId`. Download the four keys below. |

Use `SUMMARY`, `REVIEW` and `DOCUMENTS` as JSON responses. Use `PRODUCTS.csv` with the HTTP Request response format set to File. Keep its data binary until saved to your chosen staging destination.

The [Wait node](https://docs.n8n.io/integrations/builtin/core-nodes/n8n-nodes-base.wait/) supports delays; the loop condition and counter are your workflow's responsibility. A timeout ends collection, not the cloud run. Resume B with the same saved ID. FAILED, TIMED-OUT, ABORTED and unknown statuses stop collection; they never create a replacement run.

## Before scheduling or importing

- If A's POST times out or its response cannot be saved, mark the reservation as outcome unknown. Reconcile the existing run in Console; do not execute A again automatically.
- If B's GET fails, preserve its ID and resume collection later. Do not retry the whole start workflow.
- Compare CSV records with `SUMMARY.delivered_rows`. Keep diagnostics alongside each export. Successful status alone does not mean a complete catalog: spending limits, page/row limits and document failures can produce partial output.
- Keep repeated supplier revisions out of A, even after a successful export. A scheduled workflow should only pass a new document revision through the reservation step.
- After verifying a real compatible supplier run in your account, configure your desired schedule or supplier-update trigger. Set an explicit budget and validated import mapping before activating downstream writes.

An unresolved document can remain in an exception queue while other supported documents complete. No sales call, consultation or manual correction service is included.

Apify references: [Run Actor](https://docs.apify.com/api/v2/actors-runs-post), [Get run](https://docs.apify.com/api/v2/actor-run-get), [Get record](https://docs.apify.com/api/v2/key-value-store-record-get). Pricing and interfaces checked September 14, 2026; review the current listing before starting paid runs.
