Compare commits

..

5 Commits

Author SHA1 Message Date
f3nris bf273b959a feat(cortex-xdr): watermark incidents on modification_time, not creation
An XDR incident is not finished when it is created. Alerts keep joining it, an
analyst changes its status, its severity is raised. A creation_time watermark
fetches it once, the watermark moves past it, and nothing that happens
afterwards ever reaches Riposte — which is precisely the content the full fetch
exists to bring in.

modified_after filters and sorts on modification_time instead, so an incident
comes back on every change and dedup on incident_id turns the second visit into
an enrichment of the incident already there. It is now what the ingest hint
prefills; created_after stays for a one-shot backfill.

Worth knowing about that enrichment: it merges context and can fill a detection
anchor that was missing, but it does not restate the incident's severity or
status. An incident XDR later raises to critical stays at the severity it was
ingested with.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 23:50:24 +02:00
f3nris ba68d19e51 refactor(cortex-xdr): name the full fetch get-incidents-full (v1.3.1)
"fetch" said nothing next to a list of commands that all start with "get". The
command sits beside cortex-xdr-get-incidents in the picker, and the only thing
an operator needs to read there is which of the two carries everything — so the
name says it: get-incidents-full.

The id moves with it (fetch_incidents -> get_incidents_full), since the script
and the bundled mapper are bound to a command by filename. Anyone who created a
rule against the old id in the few minutes 1.3.0 was up has to point it at the
new command; the changelog says so.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 23:47:17 +02:00
f3nris a6e1245c79 feat(cortex-xdr): fetch incidents with their alerts, not a 21-field summary
incidents/get_incidents/ answers with a summary and nothing else: 21 fields,
no hosts, no users, no MITRE, no tags, and not one of the alerts the incident
aggregates. Ingesting through it leaves an incident whose raw payload says
almost nothing about what happened — and the shipped mapper had been written
for a richer shape than the endpoint ever returns, mapping hosts[0], users[0]
and mitre_* that simply are not in that response.

incidents/get_multiple_incidents_extra_data/ returns the same incidents with
39 fields, every alert in full — 156 fields each — and the file and network
artifacts. It is what the reference client fetches through (demisto/content,
CortexXDRIR.get_multiple_incidents_extra_data), and full_alert_fields must be
set or the nested alerts come back trimmed to a handful of fields.

Records arrive as {incident, alerts, network_artifacts, file_artifacts} with
each nested block wrapped as {total_count, data}. The script flattens them, so
every expression written against get_incidents keeps working — the summary's
21 fields are a subset of these 39 — while the alerts and artifacts land beside
them as plain lists, and their total_count says when a list is a sample rather
than the whole set. incident_sources is lifted into a scalar for the same
reason severity was on the alerts side: the incident-field mapper reads dotted
paths and cannot index a list.

get_incidents stays, for cheap polling, and now says in its description what it
does and does not carry.

The mapper maps the aggregate first and the first alert last, so the alert
fills in whatever the aggregate leaves silent — including the detection anchor,
since an XDR incident's detection_time is usually null while its alerts carry
theirs. Verified against the vendor's recorded response
(test_data/get_multiple_incidents_extra_data.json): 33 of 52 entries resolve,
severity critical lands on 5, source reads "XDR Agent", and the anchor falls
through to the alert's detection timestamp.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 23:41:27 +02:00
f3nris d06f8ea413 fix(cortex-xdr): the alerts endpoint never answered — wrong request dialect
get_alerts posted the incidents body, {filters, search_from, search_to, sort},
to alerts/get_alerts_by_filter_data/. That endpoint serves the alerts GRID and
speaks another dialect entirely, so every call — before this branch as much as
after it — came back a bare HTTP 500 with no hint as to why.

Shape taken from the reference client (demisto/content, Packs/ApiModules/
Scripts/CoreIRApiModule, get_alerts_by_filter_command):

  request_data.filter_data = {
    sort:   [{FIELD, ORDER}],          # a list, uppercase keys
    paging: {from, to},                # not search_from/search_to
    filter: {AND: [{SEARCH_FIELD, SEARCH_TYPE, SEARCH_VALUE}]},
  }

Severity is an enum there (SEV_040_HIGH), and several severities are OR'd, not
passed as a list. The watermark is a RANGE, since the grid has no gte operator;
its upper bound carries five minutes of slack, because our clock and the
tenant's are not the same clock. A filterless query is bounded to the last
thirty days rather than sent empty — the reference client refuses one outright,
and the grid is not meant to be asked for a whole retention.

The response needed as much work as the request. Rows arrive wrapped as
{alert_fields, incident_fields}, and mapping through that wrapper would put an
alert_fields. prefix on every expression an operator writes, so each row is
unwrapped. Two of its fields cannot be mapped as they stand: severity is the
enum code, and status.progress carries a dot INSIDE the key, which no mapping
path can express. Both are derived into severity_name and status_progress.

The mapper follows the grid's own vocabulary — internal_id, alert_name,
agent_hostname, agent_ip_addresses — and dedup moves to internal_id, since
alert_id belongs to the other API. case_id is kept as the correlation UID: it
is the join back to the incident feed.

Verified end to end against the vendor's own recorded response
(test_data/get_alerts_by_filter_results.json): 33 of 54 OCSF entries resolve on
it, severity lands on 3, the detection anchor is set, and the paging walks
0-100, 100-200, 200-250 with the truncation flag raised only when the ceiling,
not the window, ended the fetch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 23:34:26 +02:00
f3nris 17318d4225 feat(cortex-xdr): ingest alerts, not only incidents (v1.3.0)
An XDR incident is an aggregate; the SOC works the detections under it. The
alerts endpoint was already exposed as a read command, but nothing could feed
an alert rule with it — no results path, no dedup key, no watermark, no mapper,
no incident type. All five are here now, so an alert rule can be pointed at
reply.alerts the same way it is pointed at reply.incidents.

get_alerts pages past the API's 100-results-per-call ceiling: an alert feed
carries far more than a hundred detections between two polls, and whatever a
single page leaves behind is never fetched again, because the next run's
watermark has already moved past it. On an incremental fetch it also sorts
oldest first, so a window larger than the limit drops its most recent alerts —
the only ones the next poll can still see — and says so via `truncated`.

Two fixes to the incident side while in the same files:

- The severity expression compared strings, which the mapping engine cannot do
  (it reads numeric comparisons only). Every test read as false, so every
  ingested incident silently took the alert rule's default severity. The bare
  field works: Riposte maps critical/high/medium/low onto 1-5 itself.
- The incident mapper carried no `time`, so the detection anchor was missing
  and MTTD stayed empty for the whole feed. creation_time fills it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 23:26:31 +02:00
7 changed files with 506 additions and 18 deletions
@@ -0,0 +1,3 @@
name: "Cortex XDR Alert"
color: "#ef8354"
icon: "alert"
+31 -8
View File
@@ -1,8 +1,8 @@
id: cortex_xdr
name: Cortex XDR
version: 1.2.1
description: "Palo Alto Cortex XDR (public API v1) — incident ingestion + write-back, endpoint isolation/scan/delete/tagging, RTR scripts, hash block/allow lists, file quarantine/restore/retrieval, alert exclusions, external alert push (parsed/CEF), device-control violations, audits, distributions and RBAC/risk."
changelog: "1.2.1 — Connection troubleshooting: the URL is normalised to the tenant host (a pasted /public_api/v1 or console path no longer breaks the call), a non-JSON reply reports the status, content type and body instead of a bare JSON parse error, missing key/key ID is caught up front, nonce and timestamp are sent in both auth modes as the reference client does, and test_connection now probes get_incidents. 1.2.0 — Incident write-back (update_incident: status/severity/assignment/resolve comment) and external alert push (insert_parsed_alerts, insert_cef_alerts). 1.1.0 — Full command coverage: added delete/alias/tag endpoints, abort scan, original alerts, script metadata/code/snippet/exec-status, file retrieval (+details), alert exclusions, device-control violations, audits, distribution url/status/create and RBAC (users, roles, groups, risk score, risky users/hosts). 1.0.0 — Initial release: incident ingestion (get_incidents) with OCSF mapper, endpoints, isolate/unisolate, scan, hash blocklist/allowlist, quarantine/restore, run script + results, alerts retrieval, distributions and action status. Standard or Advanced API authentication."
version: 1.3.1
description: "Palo Alto Cortex XDR (public API v1) — incident and alert ingestion + write-back, endpoint isolation/scan/delete/tagging, RTR scripts, hash block/allow lists, file quarantine/restore/retrieval, alert exclusions, external alert push (parsed/CEF), device-control violations, audits, distributions and RBAC/risk."
changelog: "1.3.1 — get_incidents_full can watermark on modification_time (modified_after), which is what ingestion wants: an XDR incident keeps growing after creation, and a creation_time watermark fetches it once and never looks again, so every alert that joins it afterwards is lost. The full-incident fetch command is named cortex-xdr-get-incidents-full (id get_incidents_full), not cortex-xdr-fetch-incidents: sitting next to cortex-xdr-get-incidents in the command list, it now reads as what it is — the same call, everything included. A rule created against the old id must be pointed at the new one. 1.3.0 — Richer incident ingestion (get_incidents_full, on get_multiple_incidents_extra_data): incidents now arrive with their alerts in full and their file/network artifacts, where get_incidents only ever answered a 21-field summary carrying neither hosts, users, MITRE nor a single alert. Alert ingestion, and the alerts endpoint answers at last: get_alerts was sending the incidents dialect ({filters, search_from, search_to, sort}) to a grid endpoint that speaks request_data.filter_data (SEARCH_FIELD/SEARCH_TYPE/SEARCH_VALUE blocks, paging.from/to, sort as a list), and every call came back HTTP 500. Body rebuilt from the reference client, rows unwrapped out of their alert_fields envelope, severity code and the dotted status.progress key derived into readable fields. Alert ingestion: get_alerts is now a fetch command (results path reply.alerts, dedup on alert_id, incremental on source_insert_ts) with a bundled OCSF mapper and a Cortex XDR Alert incident type, so detections can be ingested alongside — or instead of — incidents. The incident mapper is fixed on the way past: its severity expression compared strings, which the mapping engine cannot do, so every ingested incident silently took the rule's default severity; it also now carries a detection anchor so MTTD is measurable. It pages past the API's 100-results-per-call ceiling, and sorts oldest-first on an incremental fetch so a truncated window drops the alerts the next poll can still see. 1.2.1 — Connection troubleshooting: the URL is normalised to the tenant host (a pasted /public_api/v1 or console path no longer breaks the call), a non-JSON reply reports the status, content type and body instead of a bare JSON parse error, missing key/key ID is caught up front, nonce and timestamp are sent in both auth modes as the reference client does, and test_connection now probes get_incidents. 1.2.0 — Incident write-back (update_incident: status/severity/assignment/resolve comment) and external alert push (insert_parsed_alerts, insert_cef_alerts). 1.1.0 — Full command coverage: added delete/alias/tag endpoints, abort scan, original alerts, script metadata/code/snippet/exec-status, file retrieval (+details), alert exclusions, device-control violations, audits, distribution url/status/create and RBAC (users, roles, groups, risk score, risky users/hosts). 1.0.0 — Initial release: incident ingestion (get_incidents) with OCSF mapper, endpoints, isolate/unisolate, scan, hash blocklist/allowlist, quarantine/restore, run script + results, alerts retrieval, distributions and action status. Standard or Advanced API authentication."
category: endpoint
# Per-instance configuration. The base URL is the tenant API root, e.g.
@@ -44,7 +44,7 @@ commands:
# ── Ingestion ─────────────────────────────────────────────────────────────
- id: get_incidents
name: cortex-xdr-get-incidents
description: "Fetch Cortex XDR incidents for ingestion. Returns {reply:{incidents:[...]}}; use reply.incidents as the alert rule results path."
description: "List Cortex XDR incidents as a 21-field summary (no hosts, no users, no MITRE, no alerts). Cheap to poll, but for ingestion prefer cortex-xdr-get-incidents-full, which returns the same incidents with their alerts and artifacts. Returns {reply:{incidents:[...]}}."
risk: read
inputs_schema:
properties:
@@ -57,6 +57,23 @@ commands:
results_path: reply.incidents
dedup_key: incident_id
incremental_field: created_after
- id: get_incidents_full
name: cortex-xdr-get-incidents-full
description: "Fetch incidents WITH their alerts and artifacts (get_multiple_incidents_extra_data) — the ingestion command to prefer. get_incidents answers with a 21-field summary carrying no hosts, no users, no MITRE and none of the alerts; this one returns 39 incident fields, every alert in full (156 fields each) and the file/network artifacts. Records are flattened, so mapping expressions written against get_incidents keep working and alerts[], file_artifacts[], network_artifacts[] sit beside them. Returns {reply:{incidents:[...]}}."
risk: read
inputs_schema:
properties:
status: { type: string, description: "Comma-separated statuses to keep (new, under_investigation, resolved_threat_handled…)" }
created_after: { type: string, description: "Lower bound on creation_time, ISO8601 or epoch ms. Watermarking on this fetches each incident once and never revisits it — alerts joining it later never arrive." }
modified_after: { type: string, description: "Lower bound on modification_time, ISO8601 or epoch ms. The watermark to prefer for ingestion: an incident comes back whenever it changes, and dedup on incident_id turns the second visit into an enrichment." }
limit: { type: number, description: "Maximum incidents to fetch (default 50, paged 50 at a time). A full incident weighs a few KB and up to a few hundred with its alerts, so raise this knowingly." }
exclude_artifacts: { type: boolean, description: "Drop the file and network artifact blocks, keeping the alerts (lighter payload)" }
required: []
outputs_schema: { properties: {} }
ingest:
results_path: reply.incidents
dedup_key: incident_id
incremental_field: modified_after
- id: get_incident_extra_data
name: cortex-xdr-get-incident-extra-data
description: "Get full incident data including its alerts and network artifacts by incident ID."
@@ -83,15 +100,21 @@ commands:
outputs_schema: { properties: {} }
- id: get_alerts
name: cortex-xdr-get-alerts
description: "Retrieve alerts using a custom filter (get_alerts_by_filter_data). Returns rich alert objects."
description: "Fetch Cortex XDR alerts for ingestion (get_alerts_by_filter_data). Returns {reply:{alerts:[...]}}; use reply.alerts as the alert rule results path. Each row is unwrapped out of the API's alert_fields envelope and carries a readable severity_name and status_progress, so alerts-grid field names (internal_id, alert_name, agent_hostname) are what mapping expressions see. Alerts are the detection layer under incidents: ingest them alongside get_incidents when the SOC works detections, not only aggregates."
risk: read
inputs_schema:
properties:
severity: { type: string, description: "Comma-separated severities (low, medium, high, critical)" }
created_after: { type: string, description: "Lower bound on alert source_insert_ts, epoch ms" }
limit: { type: number, description: "Maximum alerts to fetch (default 100)" }
severity: { type: string, description: "Comma-separated severities (informational, low, medium, high, critical)" }
created_after: { type: string, description: "Lower bound on alert source_insert_ts, ISO8601 or epoch ms (incremental fetch watermark)" }
limit: { type: number, description: "Maximum alerts to fetch (default 100). The API serves 100 per call at most; above that the script pages until the limit is reached." }
# Left unfiltered, the call is bounded to the last 30 days: the alerts
# grid is not meant to be asked for a tenant's whole retention.
required: []
outputs_schema: { properties: {} }
ingest:
results_path: reply.alerts
dedup_key: internal_id
incremental_field: created_after
- id: insert_parsed_alerts
name: cortex-xdr-insert-parsed-alerts
description: "Push external alerts (parsed JSON objects) into Cortex XDR for correlation."
@@ -0,0 +1,87 @@
name: "Cortex XDR Alerts → OCSF"
description: "Maps one Cortex XDR alert (alerts/get_alerts_by_filter_data/, results_path = reply.alerts) to OCSF Detection Finding fields. Field names are the alerts-grid ones (internal_id, alert_name, agent_hostname…), not the incident ones; the script unwraps the API's alert_fields envelope and derives severity_name and status_progress, which the raw payload cannot express. Alerts are the detection layer under incidents: a tenant ingesting both feeds holds each detection twice, once inside an aggregate and once on its own."
field_mappings:
title: "alert_name"
description: "alert_description"
# severity_name, not severity: the API sends an enum code (SEV_040_HIGH) that
# no severity scale can read, so the script carries the plain name alongside it.
severity: "severity_name"
# Which sensor fired: "XDR Agent", "PAN NGFW", "XDR Analytics"…
source: "alert_source"
# results_path = reply.alerts; source_path is JSONata over ONE alert object.
# Paths absent from a given alert are skipped at ingestion, so entries for fields
# a tenant never emits are safe. Where two entries target the same OCSF field,
# the LAST non-empty one wins — that is how the events[] fallbacks are ordered.
ocsf:
# ── Finding ───────────────────────────────────────────────────────
- { source_path: "internal_id", ocsf_field: "finding_info.uid" }
- { source_path: "external_id", ocsf_field: "finding_info.uid_alt" }
- { source_path: "alert_name", ocsf_field: "finding_info.title" }
- { source_path: "alert_description", ocsf_field: "finding_info.desc" }
- { source_path: "source_insert_ts", ocsf_field: "finding_info.created_time" }
- { source_path: "local_insert_ts", ocsf_field: "finding_info.modified_time" }
- { source_path: "alert_category", ocsf_field: "finding_info.analytic.category" }
- { source_path: "alert_name", ocsf_field: "finding_info.analytic.name" }
- { source_path: "matching_service_rule_id", ocsf_field: "finding_info.analytic.uid" }
# ── Detection time: the MTTD anchor ───────────────────────────────
# `time` is what Riposte measures detection-to-ingestion against. The grid
# exposes when the tenant took the alert in (source_insert_ts); a sensor-side
# detection timestamp, when the tenant sends one, is the better anchor and
# comes last so it wins.
- { source_path: "source_insert_ts", ocsf_field: "time" }
- { source_path: "detection_timestamp", ocsf_field: "time" }
# ── Alert state ───────────────────────────────────────────────────
- { source_path: "severity_name", ocsf_field: "severity" }
- { source_path: "alert_domain", ocsf_field: "activity_name" }
- { source_path: "alert_action_status", ocsf_field: "action" }
- { source_path: "status_progress", ocsf_field: "status" }
- { source_path: "matching_status", ocsf_field: "status_detail" }
- { source_path: "events_length", ocsf_field: "count" }
# The XDR case this alert was folded into — the join back to the incident feed.
- { source_path: "case_id", ocsf_field: "metadata.correlation_uid" }
# ── Product identity ──────────────────────────────────────────────
- { source_path: "'Cortex XDR'", ocsf_field: "metadata.product.name" }
- { source_path: "'Palo Alto Networks'", ocsf_field: "metadata.product.vendor_name" }
- { source_path: "alert_source", ocsf_field: "metadata.log_source" }
# ── MITRE ATT&CK ──────────────────────────────────────────────────
# Both fields arrive as a list on most tenants and as a bare string on some;
# [0] reads the first element either way.
- { source_path: "mitre_tactic_id_and_name[0]", ocsf_field: "attacks.tactic.name" }
- { source_path: "mitre_technique_id_and_name[0]", ocsf_field: "attacks.technique.name" }
# ── Affected endpoint ─────────────────────────────────────────────
- { source_path: "agent_hostname", ocsf_field: "device.hostname" }
- { source_path: "agent_ip_addresses[0]", ocsf_field: "device.ip" }
- { source_path: "agent_id", ocsf_field: "device.uid" }
- { source_path: "agent_os_type", ocsf_field: "device.os.type" }
# Mirrored onto src_endpoint so routers and pre-processing rules written for
# the incident feed (which maps hosts there) match alerts unchanged.
- { source_path: "agent_hostname", ocsf_field: "src_endpoint.hostname" }
- { source_path: "agent_ip_addresses[0]", ocsf_field: "src_endpoint.ip" }
- { source_path: "actor_effective_username", ocsf_field: "user.name" }
# ── What actually happened ────────────────────────────────────────
# Grid columns first, then the same detail read off the first event, which is
# where a tenant that does not flatten these columns puts them.
- { source_path: "actor_process_image_name", ocsf_field: "process.name" }
- { source_path: "actor_process_command_line", ocsf_field: "process.cmd_line" }
- { source_path: "actor_process_image_sha256", ocsf_field: "process.file.hashes.sha256" }
- { source_path: "causality_actor_process_command_line", ocsf_field: "process.parent_process.cmd_line" }
- { source_path: "action_file_path", ocsf_field: "file.path" }
- { source_path: "action_file_sha256", ocsf_field: "file.hashes.sha256" }
- { source_path: "action_file_md5", ocsf_field: "file.hashes.md5" }
- { source_path: "action_registry_key_name", ocsf_field: "reg_key.path" }
- { source_path: "action_registry_data", ocsf_field: "reg_value.data" }
- { source_path: "action_local_ip", ocsf_field: "src_endpoint.ip" }
- { source_path: "action_local_port", ocsf_field: "src_endpoint.port" }
- { source_path: "action_remote_ip", ocsf_field: "dst_endpoint.ip" }
- { source_path: "action_remote_port", ocsf_field: "dst_endpoint.port" }
- { source_path: "dst_action_external_hostname", ocsf_field: "dst_endpoint.hostname" }
- { source_path: "events[0].actor_process_image_name", ocsf_field: "process.name" }
- { source_path: "events[0].actor_process_command_line", ocsf_field: "process.cmd_line" }
- { source_path: "events[0].actor_process_image_path", ocsf_field: "process.path" }
- { source_path: "events[0].actor_process_image_sha256", ocsf_field: "process.file.hashes.sha256" }
- { source_path: "events[0].causality_actor_process_image_name", ocsf_field: "process.parent_process.name" }
- { source_path: "events[0].action_file_path", ocsf_field: "file.path" }
- { source_path: "events[0].action_file_sha256", ocsf_field: "file.hashes.sha256" }
- { source_path: "events[0].action_remote_ip", ocsf_field: "dst_endpoint.ip" }
- { source_path: "events[0].action_remote_port", ocsf_field: "dst_endpoint.port" }
- { source_path: "events[0].action_external_hostname", ocsf_field: "dst_endpoint.hostname" }
@@ -2,7 +2,10 @@ name: "Cortex XDR Incidents → OCSF"
description: "Maps a Cortex XDR incident (incidents/get_incidents/, results_path = reply.incidents) to OCSF finding fields. Incidents are aggregates; use get_incident_extra_data for per-alert detail."
field_mappings:
title: "incident_name"
severity: "severity = 'critical' ? 5 : (severity = 'high' ? 4 : (severity = 'medium' ? 3 : 2))"
# The raw string, not a ternary: the mapping engine compares numbers only, so
# every string test read as false and every incident landed on the rule's
# default severity. Riposte reads critical/high/medium/low onto 1-5 itself.
severity: "severity"
description: "description"
# results_path = reply.incidents; source_path is JSONata over ONE incident object.
# Paths absent from a given incident are skipped at ingestion, so extra entries are safe.
@@ -15,6 +18,9 @@ ocsf:
- { source_path: "modification_time", ocsf_field: "finding_info.modified_time" }
- { source_path: "xdr_url", ocsf_field: "finding_info.src_url" }
- { source_path: "status", ocsf_field: "status" }
# `time` is the MTTD anchor — when XDR opened the incident, as opposed to when
# Riposte ingested it. Without it the detection delay column stays empty.
- { source_path: "creation_time", ocsf_field: "time" }
- { source_path: "alert_count", ocsf_field: "count" }
# ── MITRE ATT&CK (first aggregated tactic/technique) ──────────────
- { source_path: "mitre_tactics_ids_and_names[0]", ocsf_field: "attacks.tactic.name" }
@@ -0,0 +1,83 @@
name: "Cortex XDR Incidents (full) → OCSF"
description: "Maps one Cortex XDR incident fetched with its alerts and artifacts (incidents/get_multiple_incidents_extra_data/, results_path = reply.incidents) to OCSF finding fields. The script flattens the record, so incident fields sit at the top level — every expression written against get_incidents keeps working — while alerts[], file_artifacts[] and network_artifacts[] are plain lists beside them. Incident-level values are mapped first and the first alert's equivalents last, so the alert wins wherever the aggregate says nothing."
field_mappings:
# incident_name is null on most tenants (it is only set when someone renames
# the incident), and a mapping that resolves to nothing leaves the title to
# the incident type's fallback. description is the sentence XDR itself shows.
title: "description"
description: "description"
severity: "severity"
# incident_source, not incident_sources[0]: this mapper reads dotted paths and
# cannot index a list, so the script lifts the first sensor out for it.
source: "incident_source"
# results_path = reply.incidents; source_path is JSONata over ONE flattened
# incident. Paths absent from a given incident are skipped at ingestion, so
# entries for fields a tenant never emits are safe. Where two entries target the
# same OCSF field, the LAST non-empty one wins.
ocsf:
# ── Finding ───────────────────────────────────────────────────────
- { source_path: "incident_id", ocsf_field: "finding_info.uid" }
- { source_path: "incident_name ? incident_name : description", ocsf_field: "finding_info.title" }
- { source_path: "description", ocsf_field: "finding_info.desc" }
- { source_path: "creation_time", ocsf_field: "finding_info.created_time" }
- { source_path: "modification_time", ocsf_field: "finding_info.modified_time" }
- { source_path: "xdr_url", ocsf_field: "finding_info.src_url" }
- { source_path: "alert_categories[0]", ocsf_field: "finding_info.analytic.category" }
- { source_path: "alerts[0].name", ocsf_field: "finding_info.analytic.name" }
# ── Detection time: the MTTD anchor ───────────────────────────────
# Weakest first, strongest last. detection_time is often null on an XDR
# incident, and then the first alert's own detection timestamp is the honest
# anchor; incident creation is the last resort.
- { source_path: "creation_time", ocsf_field: "time" }
- { source_path: "alerts[0].detection_timestamp", ocsf_field: "time" }
- { source_path: "detection_time", ocsf_field: "time" }
# ── Incident state ────────────────────────────────────────────────
- { source_path: "severity", ocsf_field: "severity" }
- { source_path: "status", ocsf_field: "status" }
- { source_path: "resolve_comment", ocsf_field: "status_detail" }
- { source_path: "alert_count", ocsf_field: "count" }
- { source_path: "aggregated_score", ocsf_field: "risk_score" }
- { source_path: "tags", ocsf_field: "metadata.labels" }
- { source_path: "alerts[0].action_pretty", ocsf_field: "action" }
# ── Product identity ──────────────────────────────────────────────
- { source_path: "'Cortex XDR'", ocsf_field: "metadata.product.name" }
- { source_path: "'Palo Alto Networks'", ocsf_field: "metadata.product.vendor_name" }
- { source_path: "incident_sources[0]", ocsf_field: "metadata.log_source" }
# ── MITRE ATT&CK: the aggregate, else the first alert ─────────────
- { source_path: "mitre_tactics_ids_and_names[0]", ocsf_field: "attacks.tactic.name" }
- { source_path: "mitre_techniques_ids_and_names[0]", ocsf_field: "attacks.technique.name" }
- { source_path: "alerts[0].mitre_tactic_id_and_name[0]", ocsf_field: "attacks.tactic.name" }
- { source_path: "alerts[0].mitre_technique_id_and_name[0]", ocsf_field: "attacks.technique.name" }
# ── Affected host / user ──────────────────────────────────────────
# An incident's hosts are 'hostname:agent_id' strings; an alert names them plainly.
- { source_path: "$split(hosts[0], ':')[0]", ocsf_field: "src_endpoint.hostname" }
- { source_path: "$split(hosts[0], ':')[0]", ocsf_field: "device.hostname" }
- { source_path: "alerts[0].host_name", ocsf_field: "src_endpoint.hostname" }
- { source_path: "alerts[0].host_name", ocsf_field: "device.hostname" }
- { source_path: "alerts[0].host_ip[0]", ocsf_field: "device.ip" }
- { source_path: "alerts[0].host_ip[0]", ocsf_field: "src_endpoint.ip" }
- { source_path: "alerts[0].endpoint_id", ocsf_field: "device.uid" }
- { source_path: "alerts[0].agent_os_type", ocsf_field: "device.os.type" }
- { source_path: "users[0]", ocsf_field: "user.name" }
- { source_path: "alerts[0].user_name", ocsf_field: "user.name" }
# ── What the first alert actually saw ─────────────────────────────
- { source_path: "alerts[0].actor_process_image_name", ocsf_field: "process.name" }
- { source_path: "alerts[0].actor_process_command_line", ocsf_field: "process.cmd_line" }
- { source_path: "alerts[0].actor_process_image_path", ocsf_field: "process.path" }
- { source_path: "alerts[0].actor_process_image_sha256", ocsf_field: "process.file.hashes.sha256" }
- { source_path: "alerts[0].causality_actor_process_image_name", ocsf_field: "process.parent_process.name" }
- { source_path: "alerts[0].action_file_path", ocsf_field: "file.path" }
- { source_path: "alerts[0].action_file_name", ocsf_field: "file.name" }
- { source_path: "alerts[0].action_file_sha256", ocsf_field: "file.hashes.sha256" }
- { source_path: "alerts[0].action_file_md5", ocsf_field: "file.hashes.md5" }
- { source_path: "alerts[0].action_remote_ip", ocsf_field: "dst_endpoint.ip" }
- { source_path: "alerts[0].action_remote_port", ocsf_field: "dst_endpoint.port" }
- { source_path: "alerts[0].action_external_hostname", ocsf_field: "dst_endpoint.hostname" }
# ── The artifact the incident is really about ─────────────────────
# Last, because a file artifact is the incident's verdict on the file, where
# the alert only reports what one detection touched.
- { source_path: "file_artifacts[0].file_name", ocsf_field: "file.name" }
- { source_path: "file_artifacts[0].file_sha256", ocsf_field: "file.hashes.sha256" }
- { source_path: "file_artifacts[0].file_wildfire_verdict", ocsf_field: "malware.classifications" }
- { source_path: "network_artifacts[0].network_remote_ip", ocsf_field: "dst_endpoint.ip" }
- { source_path: "network_artifacts[0].network_domain", ocsf_field: "dst_endpoint.hostname" }
+111 -9
View File
@@ -76,19 +76,121 @@ def to_ms(v):
return None
# This endpoint speaks the alerts-grid dialect, NOT the incidents one: a body of
# {filters, search_from, search_to, sort} — what incidents/get_incidents/ takes —
# is answered with a bare HTTP 500. It wants request_data.filter_data with
# SEARCH_FIELD/SEARCH_TYPE/SEARCH_VALUE blocks, paging.from/to and a sort LIST.
# Shape taken from the reference client (demisto/content,
# Packs/ApiModules/Scripts/CoreIRApiModule — get_alerts_by_filter_command).
PAGE = 100
# Severity travels as an enum code both ways. Riposte reads plain names onto its
# 1-5 scale, so alerts carry `severity_name` alongside the raw code.
SEVERITY_CODE_TO_NAME = {
"SEV_010_INFO": "informational",
"SEV_020_LOW": "low",
"SEV_030_MEDIUM": "medium",
"SEV_040_HIGH": "high",
"SEV_050_CRITICAL": "critical",
}
SEVERITY_NAME_TO_CODE = dict((v, k) for k, v in SEVERITY_CODE_TO_NAME.items())
SEVERITY_NAME_TO_CODE["info"] = "SEV_010_INFO"
# Our clock and the tenant's are not the same clock. A range that ends exactly
# now silently drops alerts the tenant stamped a few seconds ahead of us.
SKEW_MS = 5 * 60 * 1000
# Window applied when the caller passes no filter at all — see main().
DEFAULT_LOOKBACK_MS = 30 * 24 * 60 * 60 * 1000
def severity_block(value):
"""One EQ block per severity, OR'd together (the reference client's array rule)."""
blocks = [
{"SEARCH_FIELD": "severity", "SEARCH_TYPE": "EQ",
"SEARCH_VALUE": SEVERITY_NAME_TO_CODE.get(s.lower(), s.upper())}
for s in csv(value)
]
if not blocks:
return None
return blocks[0] if len(blocks) == 1 else {"OR": blocks}
def flatten(item):
"""One grid row -> one flat alert.
The API wraps every row as {alert_fields, incident_fields}. Mapping through
that wrapper would put an `alert_fields.` prefix on every expression an
operator writes, so the row is unwrapped here and the two fields Riposte
cannot express are derived: `status.progress` carries a dot INSIDE the key
(unusable as a mapping path) and severity is an enum code.
"""
fields = item.get("alert_fields")
alert = dict(fields) if isinstance(fields, dict) else dict(item)
alert.pop("incident_fields", None)
if "status.progress" in alert:
alert["status_progress"] = alert.pop("status.progress")
name = SEVERITY_CODE_TO_NAME.get(alert.get("severity"))
if name:
alert["severity_name"] = name
incident = item.get("incident_fields")
if isinstance(incident, dict):
alert["incident_fields"] = incident
return alert
def main():
inputs = json.loads(os.environ.get("INTEGRATION_INPUTS", "{}"))
limit = int(inputs.get("limit") or 100)
filters = []
if inputs.get("severity"):
filters.append({"field": "severity", "operator": "in", "value": csv(inputs["severity"])})
limit = max(1, int(inputs.get("limit") or 100))
conditions = []
sev = severity_block(inputs.get("severity"))
if sev:
conditions.append(sev)
created_ms = to_ms(inputs.get("created_after"))
if created_ms is not None:
filters.append({"field": "source_insert_ts", "operator": "gte", "value": created_ms})
rd = {"search_from": 0, "search_to": limit, "sort": {"field": "source_insert_ts", "keyword": "desc"}}
if filters:
rd["filters"] = filters
print(json.dumps(post("/alerts/get_alerts_by_filter_data/", rd)))
conditions.append({
"SEARCH_FIELD": "source_insert_ts",
"SEARCH_TYPE": "RANGE",
"SEARCH_VALUE": {"from": created_ms, "to": int(time.time() * 1000) + SKEW_MS},
})
if not conditions:
# The reference client refuses a filterless query outright, and an
# unbounded scan of the whole alerts grid is not what the API is for.
# A recent window is a better default than an empty filter the tenant
# may well answer with a 500.
now_ms = int(time.time() * 1000)
conditions.append({
"SEARCH_FIELD": "source_insert_ts",
"SEARCH_TYPE": "RANGE",
"SEARCH_VALUE": {"from": now_ms - DEFAULT_LOOKBACK_MS, "to": now_ms + SKEW_MS},
})
# Oldest first on an incremental fetch, so that a window holding more alerts
# than `limit` drops its most RECENT ones — the only ones the next poll can
# still see. Newest first otherwise, which is what an operator running the
# command by hand is asking for.
order = "ASC" if created_ms is not None else "DESC"
alerts, truncated = [], False
while len(alerts) < limit:
rd = {"filter_data": {
"sort": [{"FIELD": "source_insert_ts", "ORDER": order}],
"paging": {"from": len(alerts), "to": min(len(alerts) + PAGE, limit)},
"filter": {"AND": conditions},
}}
reply = (post("/alerts/get_alerts_by_filter_data/", rd) or {}).get("reply") or {}
page = reply.get("alerts") or []
alerts.extend(flatten(a) for a in page)
if len(page) < PAGE:
break
# Stopped on the ceiling rather than on an exhausted window: whatever is
# left is not coming back on the next poll, and a silent cap reads like
# a quiet feed.
truncated = len(alerts) >= limit
out = {"result_count": len(alerts), "alerts": alerts}
if truncated:
out["truncated"] = True
print(json.dumps({"reply": out}))
try:
@@ -0,0 +1,184 @@
import json, os, sys, time, hashlib, secrets, string, urllib.request, urllib.error
from datetime import datetime
def _client():
s = json.loads(os.environ.get("INTEGRATION_SECRETS") or "{}")
raw = str(s.get("url") or "").strip().rstrip("/")
if not raw:
raise ValueError("no url configured — paste the tenant API URL (Cortex XDR > Settings > Configurations > API Keys > Copy URL)")
if "://" not in raw:
raw = "https://" + raw
# The tenant URL is a bare host. Drop whatever was pasted after it (a stray
# /public_api/v1, a console path) so the API root is built exactly once.
scheme, _, rest = raw.partition("://")
base = scheme + "://" + rest.split("/")[0] + "/public_api/v1"
key = s.get("api_key", "")
kid = str(s.get("api_key_id", ""))
if not key or not kid:
raise ValueError("api_key and api_key_id are both required")
# Nonce and timestamp ride along in both modes, as the reference client does.
# A standard key travels as-is; an advanced one as sha256(key + nonce + ts).
nonce = "".join(secrets.choice(string.ascii_letters + string.digits) for _ in range(64))
ts = str(int(time.time()) * 1000)
headers = {
"x-xdr-auth-id": kid,
"x-xdr-nonce": nonce,
"x-xdr-timestamp": ts,
"Content-Type": "application/json",
"Accept": "application/json",
}
if str(s.get("auth_type") or "standard").lower() == "advanced":
headers["Authorization"] = hashlib.sha256((key + nonce + ts).encode("utf-8")).hexdigest()
else:
headers["Authorization"] = key
return base, headers
def _not_json(r, raw):
"""A 2xx that is not JSON means we are not talking to the XDR API at all."""
ctype = (r.headers.get("Content-Type") or "unknown").split(";")[0].strip()
head = raw[:160].decode("utf-8", "replace").replace("\n", " ").strip()
return (
"expected JSON from " + r.geturl() + ", got " + ctype + " (HTTP " + str(r.status) + "): " + head
+ " — check the configured url is the tenant API host"
+ " (https://api-<tenant>.xdr.<region>.paloaltonetworks.com), not the console URL"
)
def post(path, request_data):
base, headers = _client()
data = json.dumps({"request_data": request_data}).encode("utf-8")
req = urllib.request.Request(base + path, data=data, headers=headers, method="POST")
with urllib.request.urlopen(req, timeout=90) as r:
raw = r.read()
if not raw:
return {}
try:
return json.loads(raw)
except ValueError:
raise ValueError(_not_json(r, raw))
def to_ms(v):
if v in (None, ""):
return None
s = str(v)
if s.isdigit():
return int(s)
try:
return int(datetime.fromisoformat(s.replace("Z", "+00:00")).timestamp() * 1000)
except Exception:
return None
def csv(v):
return [x.strip() for x in str(v or "").split(",") if x.strip()]
# incidents/get_incidents/ answers with a 21-field summary: no hosts, no users,
# no MITRE, and above all not one of the alerts the incident aggregates. This
# endpoint returns the same incident with 39 fields, its alerts in full (156
# fields each) and its file/network artifacts — which is why the reference
# client fetches through it and not through get_incidents (demisto/content,
# CortexXDRIR.get_multiple_incidents_extra_data).
PAGE = 50
# Artifacts are dropped by name, not by omission — the API only understands
# being told which blocks to leave out.
ARTIFACT_BLOCKS = ["network_artifacts", "file_artifacts"]
def flatten(item):
"""One record -> one incident.
Records arrive as {incident, alerts, network_artifacts, file_artifacts},
each nested block wrapped as {total_count, data}. Flattening the incident to
the top level keeps every expression written against get_incidents working
unchanged — the summary's 21 fields are a subset of these 39 — while the
alerts and artifacts land beside them as plain lists.
"""
incident = dict(item.get("incident") or {})
for key in ("alerts", "network_artifacts", "file_artifacts"):
block = item.get(key)
if not isinstance(block, dict):
continue
incident[key] = block.get("data") or []
if block.get("total_count") is not None:
# The tenant caps alerts per incident (50 by default), so the count
# says when the list is a sample rather than the whole set.
incident[key + "_total_count"] = block["total_count"]
# The producing sensor is a list here, and the incident-field mapper reads
# dotted paths only — no array indexing — so the first source is lifted out
# for it. The list itself stays, for expressions that can index.
sources = incident.get("incident_sources")
if isinstance(sources, list) and sources:
incident["incident_source"] = sources[0]
return incident
def main():
inputs = json.loads(os.environ.get("INTEGRATION_INPUTS", "{}"))
limit = max(1, int(inputs.get("limit") or 50))
filters = []
if inputs.get("status"):
statuses = csv(inputs["status"])
filters.append({"field": "status", "operator": "in", "value": statuses})
created_ms = to_ms(inputs.get("created_after"))
if created_ms is not None:
filters.append({"field": "creation_time", "operator": "gte", "value": created_ms})
# An XDR incident keeps growing after it is created: alerts join it, an
# analyst changes its status. Watermarking on creation_time fetches it once
# and never looks again, so everything that happened afterwards is lost.
# Watermarking on modification_time brings it back on every change, where
# dedup on incident_id turns the second visit into an enrichment.
modified_ms = to_ms(inputs.get("modified_after"))
if modified_ms is not None:
filters.append({"field": "modification_time", "operator": "gte", "value": modified_ms})
incremental = created_ms is not None or modified_ms is not None
# Oldest first on an incremental fetch, so that a window holding more
# incidents than `limit` drops its most RECENT ones — the only ones the next
# poll can still see. Newest first otherwise, for a hand-run command.
sort_field = "modification_time" if modified_ms is not None else "creation_time"
keyword = "asc" if incremental else "desc"
exclude = str(inputs.get("exclude_artifacts") or "").lower() in ("1", "true", "yes")
incidents, total = [], None
while len(incidents) < limit:
rd = {
"search_from": len(incidents),
"search_to": min(len(incidents) + PAGE, limit),
"sort": {"field": sort_field, "keyword": keyword},
# Without this the nested alerts come back trimmed to a handful of
# fields — the very thing this command exists to avoid.
"full_alert_fields": True,
}
if filters:
rd["filters"] = filters
if exclude:
rd["fields_to_exclude"] = ARTIFACT_BLOCKS
reply = (post("/incidents/get_multiple_incidents_extra_data/", rd) or {}).get("reply") or {}
page = reply.get("incidents") or []
if total is None:
total = reply.get("total_count")
incidents.extend(flatten(i) for i in page)
if len(page) < PAGE:
break
out = {"result_count": len(incidents), "incidents": incidents}
if total is not None:
out["total_count"] = total
# Say it when the window held more than the limit: those incidents are
# not coming back on the next poll.
out["truncated"] = total > len(incidents)
print(json.dumps({"reply": out}))
try:
main()
except urllib.error.HTTPError as e:
print(json.dumps({"error": "HTTP " + str(e.code), "detail": e.read().decode("utf-8", "replace")}))
sys.exit(1)
except Exception as e:
print(json.dumps({"error": str(e)}))
sys.exit(1)