Internal API
Onboarding
21 endpoints.
/api/v1/internal/workflow-engine/onboarding/crawl-attempts/{attempt_id}/assetsIngest Assets
Checkpointed batch of downloaded+stored assets (metadata + GCS blob name + DOM signals). Idempotent by (crawl_attempt_id, asset_url).
Parameters
| Name | In | Type | Description |
|---|---|---|---|
attempt_idrequired | path | string (uuid) | |
x-internal-api-key | header | string | null |
Request bodyrequired
application/jsonIngestAssetsBody| Field | Type | Description | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
assets | IngestAssetItem[] | IngestAssetItem[] fields
|
Responses
200Successful Responseapplication/jsonSuccessResponse_IngestAssetsData_| Field | Type | Description | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
request_idrequired | string | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
success | true | default true | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
datarequired | IngestAssetsData | IngestAssetsData fields
| |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
metadata | ResponseMetadata | null | ResponseMetadata fields
| |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
meta | ResponseMeta | null | ResponseMeta fields
|
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/crawl-attempts/{attempt_id}/classifyClassify Crawl Pages
Label every crawled page and return the fan-out routing table (W3).
One cheap batched LLM call over compact page descriptors. The response is the engine's whole work list: ``batches`` (one extract-domain call each), ``blog_post_page_ids`` (one extract-blog-post call each) and ``other_page_ids``. Every page appears in at least one of the three — a page the classifier missed is assigned ``other`` and reported in ``unrouted_page_ids``, never dropped.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
attempt_idrequired | path | string (uuid) | |
x-internal-api-key | header | string | null |
Responses
200Successful Responseapplication/jsonSuccessResponse_ClassifyPagesData_| Field | Type | Description | |||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
request_idrequired | string | ||||||||||||||||||||||
success | true | default true | |||||||||||||||||||||
datarequired | ClassifyPagesData | ClassifyPagesData fields
| |||||||||||||||||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | |||||||||||||||||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/crawl-attempts/{attempt_id}/completeComplete Crawl
Finalize the crawl attempt. On ``completed`` runs media promotion inline (scoped to this attempt) so scraped images enter the photo library.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
attempt_idrequired | path | string (uuid) | |
x-internal-api-key | header | string | null |
Request bodyrequired
application/jsonCompleteCrawlBody| Field | Type | Description | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
statusrequired | string | ||||||||||||||||
counters | CompleteCrawlCounters | null | CompleteCrawlCounters fields
| |||||||||||||||
summary_json | object | null | ||||||||||||||||
error_message | string | null |
Responses
200Successful Responseapplication/jsonSuccessResponse_CompleteCrawlData_| Field | Type | Description | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
request_idrequired | string | |||||||||||||
success | true | default true | ||||||||||||
datarequired | CompleteCrawlData | CompleteCrawlData fields
| ||||||||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | ||||||||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/crawl-attempts/{attempt_id}/extract-blog-postExtract Blog Post Fragment
Extract one blog post from one page (W5).
The finest fan-out grain: one call per post is what restores full article bodies. The engine treats a post that exhausts its retries as best-effort — it drops the fragment and merge's CSS fallback covers the post.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
attempt_idrequired | path | string (uuid) | |
x-internal-api-key | header | string | null |
Request bodyrequired
application/jsonExtractBlogPostBody| Field | Type | Description |
|---|---|---|
page_idrequired | string |
Responses
200Successful Responseapplication/jsonSuccessResponse_ExtractBlogPostData_| Field | Type | Description | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
request_idrequired | string | ||||||||||
success | true | default true | |||||||||
datarequired | ExtractBlogPostData | ExtractBlogPostData fields
| |||||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | |||||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/crawl-attempts/{attempt_id}/extract-domainExtract Domain Fragment
Extract one routed batch (≤5 pages) for one data domain (W4).
The batch's full page content plus the domain's current-SOR slice go into one LLM call whose whole output budget belongs to this batch. Returns a payload-shaped fragment — off-domain collections included, because an FAQ sitting on a service page must survive a classifier mislabel.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
attempt_idrequired | path | string (uuid) | |
x-internal-api-key | header | string | null |
Request bodyrequired
application/jsonExtractDomainBody| Field | Type | Description |
|---|---|---|
domainrequired | string | |
page_ids | string[] | string[] fields |
Responses
200Successful Responseapplication/jsonSuccessResponse_ExtractDomainData_| Field | Type | Description | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
request_idrequired | string | |||||||||||||
success | true | default true | ||||||||||||
datarequired | ExtractDomainData | ExtractDomainData fields
| ||||||||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | ||||||||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/crawl-attempts/{attempt_id}/extract-findingsExtract Findings
Sweep every crawled page for off-schema value → typed ``extra:*`` contexts (W6).
Additive by design: an empty crawl returns an empty list rather than failing, and malformed entries are dropped instead of costing the good ones.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
attempt_idrequired | path | string (uuid) | |
x-internal-api-key | header | string | null |
Request bodyrequired
application/jsonExtractFindingsBody| Field | Type | Description |
|---|---|---|
captured_domains | string[] | string[] fields |
Responses
200Successful Responseapplication/jsonSuccessResponse_ExtractFindingsData_| Field | Type | Description | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
request_idrequired | string | ||||||||||
success | true | default true | |||||||||
datarequired | ExtractFindingsData | ExtractFindingsData fields
| |||||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | |||||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/crawl-attempts/{attempt_id}/fetched-urlsFetched Urls
Normalized URLs already persisted for this attempt — the scrape job seeds its visited-set from this so a retry skips completed pages.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
attempt_idrequired | path | string (uuid) | |
x-internal-api-key | header | string | null |
Responses
200Successful Responseapplication/jsonSuccessResponse_FetchedUrlsData_| Field | Type | Description | ||||||
|---|---|---|---|---|---|---|---|---|
request_idrequired | string | |||||||
success | true | default true | ||||||
datarequired | FetchedUrlsData | FetchedUrlsData fields
| ||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | ||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/crawl-attempts/{attempt_id}/heartbeatHeartbeat Crawl Attempt
Proof of life from the scrape job, independent of ingest traffic.
Pages reach sor per batch and assets only after the crawl, so a job can be busy for longer than the stale window without calling ``/pages`` or ``/assets`` — most plainly a retry that re-crawls the pages a killed execution already persisted. The job beats this on a timer for the whole crawl + asset pass. No-op (``beat=false``) once the attempt is terminal, so a superseded execution cannot keep a finished run looking alive.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
attempt_idrequired | path | string (uuid) | |
x-internal-api-key | header | string | null |
Responses
200Successful Responseapplication/jsonSuccessResponse_CrawlHeartbeatData_| Field | Type | Description | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
request_idrequired | string | ||||||||||
success | true | default true | |||||||||
datarequired | CrawlHeartbeatData | CrawlHeartbeatData fields
| |||||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | |||||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/crawl-attempts/{attempt_id}/pagesList Pages
Pages of a crawl attempt, in one of two shapes (design §11).
``fields=descriptor`` (default): compact classifier descriptors — url, title, H1/H2 headings, first-500-char excerpt. Plain parsing, no LLM. ``fields=content``: full untruncated page text for the extractors, from the same source the legacy extraction path reads (parity contract).
NOTE: ``fields=content`` has NO consumer in the pipeline any more. It existed only to ship page text to the engine's extraction activities; those now run inside sor (``/classify``, ``/extract-domain``, …) and read the pages straight from the repository. Kept for ad-hoc inspection pending a decision on removing it.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
attempt_idrequired | path | string (uuid) | |
fields | query | string | default "descriptor" |
x-internal-api-key | header | string | null |
Responses
200Successful Responseapplication/jsonSuccessResponse_Union_PageDescriptorsData__PageContentsData__| Field | Type | Description | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
request_idrequired | string | |||||||||||||
success | true | default true | ||||||||||||
datarequired | PageDescriptorsData | PageContentsData | PageDescriptorsData | PageContentsData fieldsPageDescriptorsData
PageContentsData
| ||||||||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | ||||||||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/crawl-attempts/{attempt_id}/pagesIngest Pages
Checkpointed batch of crawled pages. Idempotent via the (crawl_attempt_id, normalized_url) unique constraint.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
attempt_idrequired | path | string (uuid) | |
x-internal-api-key | header | string | null |
Request bodyrequired
application/jsonIngestPagesBody| Field | Type | Description | ||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
pages | IngestPageItem[] | IngestPageItem[] fields
|
Responses
200Successful Responseapplication/jsonSuccessResponse_IngestPagesData_| Field | Type | Description | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
request_idrequired | string | ||||||||||
success | true | default true | |||||||||
datarequired | IngestPagesData | IngestPagesData fields
| |||||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | |||||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/photos/tagsUpdate Photo Tags
Merge vision categories into ``photo_metadata.tags`` (S5 / design §13 [3]).
Set-union per photo — tags an admin or an earlier pass added are never dropped. Unknown (or deleted) photo_ids are skipped and reported rather than 404-ing the batch: the categoriser's write must land for every photo that still exists.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
x-internal-api-key | header | string | null |
Request bodyrequired
application/jsonPhotoTagsBody| Field | Type | Description | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
updates | PhotoTagUpdateItem[] | PhotoTagUpdateItem[] fields
|
Responses
200Successful Responseapplication/jsonSuccessResponse_PhotoTagsData_| Field | Type | Description | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
request_idrequired | string | ||||||||||
success | true | default true | |||||||||
datarequired | PhotoTagsData | PhotoTagsData fields
| |||||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | |||||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/runs/{run_id}/assembleAssemble Run Payload
Assemble every fan-out fragment into one payload and merge it (W6).
Concatenation only — dedup is merge planning's job. Routes through the same ``apply_assembled_merge`` as ``/runs/{run_id}/merge`` so the fan-out never grows a second promotion path; provenance is stamped ``website scrape fanout``. Response data is the merge counts summary.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
run_idrequired | path | string (uuid) | |
x-internal-api-key | header | string | null |
Request bodyrequired
application/jsonAssemblePayloadBody| Field | Type | Description |
|---|---|---|
fragments | object[] | object[] fields |
contexts | object[] | object[] fields |
extract_attempt_id | string (uuid) | null |
Responses
200Successful Responseapplication/jsonSuccessResponse_dict_str__Any__| Field | Type | Description |
|---|---|---|
request_idrequired | string | |
success | true | default true |
datarequired | object | |
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. |
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/runs/{run_id}/categorise-imagesCategorise Run Images
Categorise the run's promoted images and auto-assign the logo slots (W7).
One deterministic logo shortlist + one capped multimodal LLM call + two writes (tags merged, slots filled only-if-empty). The engine is a thin orchestrator over this call: all the vision logic and every guardrail — the shortlist-only logo pick, the >=180px favicon floor, the never-clobber slot rule — live in sor, next to the photo data they judge.
Takes no body. Same 200-with-``status`` contract as ``/stages/{stage}``: a run with no photos comes back ``skipped`` (the LLM is never called) and a domain failure comes back ``failed`` for the engine to interpret, while infra faults — and an LLM reply that should simply be re-asked — 5xx so Temporal retries them.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
run_idrequired | path | string (uuid) | |
x-internal-api-key | header | string | null |
Responses
200Successful Responseapplication/jsonSuccessResponse_CategoriseImagesData_| Field | Type | Description | ||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
request_idrequired | string | |||||||||||||||||||||||||||||||||||||
success | true | default true | ||||||||||||||||||||||||||||||||||||
datarequired | CategoriseImagesData | CategoriseImagesData fields
| ||||||||||||||||||||||||||||||||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | ||||||||||||||||||||||||||||||||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/runs/{run_id}/crawl-attemptsCreate Crawl Attempt
Open a crawl attempt for a run. Idempotent by natural key: if the run already has a *running* attempt (a retried CreateCrawlAttempt activity), it is returned instead of opening a second one. ``Idempotency-Key`` is accepted for tracing/forward-compat but the running-attempt check is the guard.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
run_idrequired | path | string (uuid) | |
Idempotency-Key | header | string | null | |
x-internal-api-key | header | string | null |
Request bodyrequired
application/jsonCreateCrawlAttemptBody| Field | Type | Description |
|---|---|---|
crawler_provider | string | default "crawl4ai" |
Responses
200Successful Responseapplication/jsonSuccessResponse_CreateCrawlAttemptData_| Field | Type | Description | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
request_idrequired | string | ||||||||||
success | true | default true | |||||||||
datarequired | CreateCrawlAttemptData | CreateCrawlAttemptData fields
| |||||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | |||||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/runs/{run_id}/enrichEnrich Run
Run WS4 enrichment for one run (design §14) — the whole pass in one call.
Deterministic quality gate over the run's business → (only when it flags something) one fact-locked LLM call → partial payload through ``apply_assembled_merge`` with ``source_kind: "enrichment"``. When the gate flags nothing and the FAQ count is already sufficient the LLM is never called and the response is ``status="skipped"``.
No request body and no idempotency key: the gate re-reads current SOR state every call, so an already-enriched business is simply no longer thin and a repeat invocation short-circuits. That is what makes this endpoint safe for the console to expose as a button later — a fixed per-run key would instead swallow a deliberate second pass.
Failure split matches ``/stages/{stage}``: a stage-domain failure returns **200** with ``status="failed"`` so the engine applies its own fatality policy (enrichment is best-effort), while infra faults — and a malformed LLM reply, which is worth re-asking — 5xx so Temporal retries.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
run_idrequired | path | string (uuid) | |
x-internal-api-key | header | string | null |
Responses
200Successful Responseapplication/jsonSuccessResponse_EnrichRunData_| Field | Type | Description | |||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
request_idrequired | string | ||||||||||||||||||||||||||||||||||
success | true | default true | |||||||||||||||||||||||||||||||||
datarequired | EnrichRunData | EnrichRunData fields
| |||||||||||||||||||||||||||||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | |||||||||||||||||||||||||||||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/runs/{run_id}/mergeMerge Run
Merge an engine-assembled extraction payload into the SOR (S4).
The single promotion boundary: the WS2 fan-out assembler and WS4 enrichment land through the same ``merge_payload`` the legacy chain uses. Typed ``contexts`` list entries (``extra:<kind>``) each become their own business_context_versions row. Response data is the merge counts summary.
Failures 5xx on purpose (unlike ``/stages/{stage}``): the merge activity's Temporal retry policy treats transport/5xx as retryable and 4xx as fatal, which is exactly the split a planner-LLM hiccup vs a bad payload needs.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
run_idrequired | path | string (uuid) | |
x-internal-api-key | header | string | null |
Request bodyrequired
application/jsonMergeRunBody| Field | Type | Description | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
payloadrequired | object | ||||||||||
provenancerequired | MergeProvenance | MergeProvenance fields
|
Responses
200Successful Responseapplication/jsonSuccessResponse_dict_str__Any__| Field | Type | Description |
|---|---|---|
request_idrequired | string | |
success | true | default true |
datarequired | object | |
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. |
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/runs/{run_id}/photosList Run Photos
Read-only compact list of the run's promoted photos (scrape-source, live) — the same projection the in-sor categoriser scores, exposed for inspection and for any caller that wants the raw inventory. See ``image_categorisation.run_photo_view`` for the shape.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
run_idrequired | path | string (uuid) | |
x-internal-api-key | header | string | null |
Responses
200Successful Responseapplication/jsonSuccessResponse_RunPhotosData_| Field | Type | Description | ||||||
|---|---|---|---|---|---|---|---|---|
request_idrequired | string | |||||||
success | true | default true | ||||||
datarequired | RunPhotosData | RunPhotosData fields
| ||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | ||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/runs/{run_id}/sor-stateGet Sor State
Read-only compact slice of the business's current SOR state for one extraction domain (design §11), so a WS2/WS4 extractor sees what already exists and its output dedups well at merge time.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
run_idrequired | path | string (uuid) | |
domainrequired | query | string | |
x-internal-api-key | header | string | null |
Responses
200Successful Responseapplication/jsonSuccessResponse_SorStateData_| Field | Type | Description | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
request_idrequired | string | ||||||||||
success | true | default true | |||||||||
datarequired | SorStateData | SorStateData fields
| |||||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | |||||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/runs/{run_id}/stage-statusRecord Stage Status
Project a workflow stage transition into ``onboarding_json.stages``.
The admin console's onboarding checklist reads the same shape whether the run was driven by the legacy in-process chain or the engine; ``executor`` and ``workflow_run_id`` are what distinguish them. The stage key is not validated against a fixed list so later workstreams (classification, enrichment) can journal new stages without a sor deploy.
The body is merged over the stage's existing entry, not substituted for it. ``/stages/{stage}`` and ``/stage-status`` are two calls, and keys the wrapper produced must survive the second one even if the engine doesn't echo them: ``finalize_website_onboarding`` finds the run by the ``website_id`` the website stage journaled, so dropping it would strand that stage at ``running`` forever.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
run_idrequired | path | string (uuid) | |
x-internal-api-key | header | string | null |
Request bodyrequired
application/jsonStageStatusBody| Field | Type | Description |
|---|---|---|
stagerequired | string | min length 1 |
statusrequired | string | |
error | string | null | |
workflow_run_id | string | null | |
outcome | object | null |
Responses
200Successful Responseapplication/jsonSuccessResponse_StageStatusData_| Field | Type | Description | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
request_idrequired | string | ||||||||||
success | true | default true | |||||||||
datarequired | StageStatusData | StageStatusData fields
| |||||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | |||||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/runs/{run_id}/stages/{stage}Run Stage
Run one onboarding stage synchronously.
V1 of the engine-owned pipeline wraps the existing sor stage implementations rather than reimplementing them, so the legacy chain and the workflow drive identical code.
A stage-level failure returns **200** with ``status="failed"`` and the error: the workflow engine owns retry and fatality policy (merge/website are fatal, company_name/industries are not), so it must see the outcome rather than a 5xx it would blindly retry. Genuine infrastructure faults still 5xx.
Deliberately does not journal — the engine is the single writer of ``onboarding_json.stages`` via ``/stage-status``.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
run_idrequired | path | string (uuid) | |
stagerequired | path | string | |
x-internal-api-key | header | string | null |
Request bodyrequired
application/jsonRunStageBody| Field | Type | Description |
|---|---|---|
extract_attempt_id | string (uuid) | null |
Responses
200Successful Responseapplication/jsonSuccessResponse_RunStageData_| Field | Type | Description | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
request_idrequired | string | ||||||||||||||||
success | true | default true | |||||||||||||||
datarequired | RunStageData | RunStageData fields
| |||||||||||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | |||||||||||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.
/api/v1/internal/workflow-engine/onboarding/runs/{run_id}/website-photos/auto-assignAuto Assign Website Photos
Fill the logo/favicon ``website_photos`` slots — only where empty.
An admin (or earlier-run) assignment always wins: a filled slot is reported ``already_set`` and never clobbered (design §13 guardrails). The website is resolved through the run's business. A photo_id that isn't a live photo of that business is skipped as ``unknown_photo`` so a bad id can't wedge a slot with a dangling reference.
Parameters
| Name | In | Type | Description |
|---|---|---|---|
run_idrequired | path | string (uuid) | |
x-internal-api-key | header | string | null |
Request bodyrequired
application/jsonAutoAssignWebsitePhotosBody| Field | Type | Description |
|---|---|---|
logo_full | string (uuid) | null | |
favicon | string (uuid) | null |
Responses
200Successful Responseapplication/jsonSuccessResponse_AutoAssignWebsitePhotosData_| Field | Type | Description | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
request_idrequired | string | ||||||||||
success | true | default true | |||||||||
datarequired | AutoAssignWebsitePhotosData | AutoAssignWebsitePhotosData fields
| |||||||||
metadata | ResponseMetadata | null | ResponseMetadata fieldsResponseMetadata, expanded above. | |||||||||
meta | ResponseMeta | null | ResponseMeta fieldsResponseMeta, expanded above. |
400Bad request401Unauthorized403Forbidden404Not found422Validation error500Internal server error503Service unavailable
Error bodies: ErrorResponse. See Errors.