Security observability · ClickHouse

Want to store Petabytes of security data for cheap?

ClickHouse keeps compressed, columnar data on inexpensive object storage and still answers investigation queries quickly. We design the pipeline, schemas and retention, move your high-volume sources over, and hand it to your team.

Talk to an engineer

SIEMs are good for audit logs. ClickHouse is good for everything else.

Keep in the SIEM

  • Alerts and triage
  • Case management and response
  • Compliance-scoped audit logs

Move to ClickHouse

  • Runtime telemetry (Tetragon, Falco)
  • Network flows (Cilium Hubble, VPC Flow Logs)
  • CloudTrail data events and DNS at full volume
  • Application, proxy and WAF logs

Detections run on ClickHouse and send high-confidence alerts to the SIEM your team already works in, so the SIEM holds what it is best at.

Splunk costing too much?

Ingest-based licensing turns every new log source into a budget argument, and teams end up dropping the data they will want during an incident. We start by measuring what each source costs you today and how often it is actually searched.

  1. Rank sources by ingest volume against how often they are queried.
  2. Move the high-volume, rarely-searched ones to ClickHouse on S3.
  3. Keep alerting in your SIEM, fed only with what needs an alert.

The result is a cost model you can compare with your current bill before you commit to any of it.

Need to keep up with AI observability? ClickHouse handles it well.

Agents leave a trail at every step: prompts, tool calls, model responses, token counts and latencies. Agent traces grow fast, and per-GB pricing grows with them. It is the kind of high-volume event data ClickHouse is built for.

  • Trace every agent and tool call, and query what an agent did and for whom.
  • Track token usage and cost by team, model and workflow.
  • Keep long retention for audit, with sensitive fields redacted before storage.

We set up the schema and the pipeline behind the OpenTelemetry collectors you may already run.

What we build

Ingestion and parsing

Per-source pipelines with parsing and field extraction, deduplication, late-arriving data handling, and replay from S3 when a parser changes.

Schema and indexing

A table design per source type, with ordering keys and skip indexes for IPs, hashes and users, chosen from the searches you run.

Search and IOC sweeps

Grep-style search across raw logs, and fast sweeps for IPs, domains and hashes across months of history.

Migrating your searches

Existing SPL, KQL or ES|QL searches and dashboards ported to ClickHouse SQL, and checked against the results in the old system.

Detections and alerting

Scheduled detections with thresholds, joins and allowlists, kept as code and alerting into your SIEM, Slack or PagerDuty.

Enrichment

GeoIP, asset and identity context, and threat intel lookups, applied at ingest or at query time.

Access and audit

SSO, role-based access, row-level policies per team, and an audit trail of who queried what.

Retention and cost

Per-source retention like the index policies you have now, hot and cold tiers on S3, and a report of storage and ingest cost by source.

Health and operations

Alerts for silent sources and ingestion lag, plus backups, upgrades and runbooks your team can follow.

Example sources we handle

  • Tetragon
  • Falco
  • Cilium Hubble
  • CloudTrail
  • VPC Flow Logs
  • GuardDuty
  • DNS logs
  • WAF logs
  • Kubernetes audit
  • Application logs

How it runs

  1. 1

    Data review

    Sources, daily volume, retention requirements and the investigations you run.

  2. 2

    Design

    Schema, ingestion architecture and a cost model against what you spend today.

  3. 3

    Build

    Working pipeline and tables loaded with your real data.

  4. 4

    Handoff

    Query library, runbooks and a walkthrough for the teams who will own it.

ClickHouse is a trademark of ClickHouse, Inc. Palm Sec is not affiliated with or endorsed by ClickHouse, Inc.