Skip to main content

Security Scanning with the Atlas Kubernetes Operator

The AtlasSecurityScan resource runs atlas security scan from the Atlas Kubernetes Operator. It scans a database on a cron schedule, and again whenever an AtlasSchema or AtlasMigration it references applies a change. It then grades any findings against a policy and stores them in an AtlasSecurityReport owned by the scan. The verdict is exposed as Kubernetes conditions, which kubectl wait and GitOps tools read directly.

The AtlasSecurityScan resource is available to Atlas Pro users with the Security Graph enabled in their plan. The Operator authenticates with a bot token. Install the Operator first, as described in Installation.

Creating a Scan​

The examples run in the app namespace, next to the application database:

kubectl config set-context --current --namespace=app

Store the database URL and the Atlas token in Secrets:

kubectl create secret generic app-db \
--from-literal=url='postgres://postgres:pass@postgres.app:5432/app?sslmode=disable'
kubectl create secret generic atlas-token --from-literal=ATLAS_TOKEN="$ATLAS_TOKEN"

The Operator connects from its own namespace, so a database running in the cluster is addressed by its namespace-qualified service name, postgres.app above. Extensions are read at the database level, so the URL does not need a search_path: a schema-bound URL reports the same extensions.

Create a scan that runs every morning at 06:00 UTC and fails the policy on a HIGH finding:

security-scan.yaml
apiVersion: db.atlasgo.io/v1alpha1
kind: AtlasSecurityScan
metadata:
name: app
spec:
urlFrom:
secretKeyRef:
name: app-db
key: url
cloud:
tokenFrom:
secretKeyRef:
name: atlas-token
key: ATLAS_TOKEN
schedule: "0 6 * * *"
timeZone: UTC
policy:
minSeverity: ELEVATED
failOn: HIGH

The first scan runs as soon as the resource is created, without waiting for the first schedule slot:

kubectl apply -f security-scan.yaml
kubectl wait atlassecurityscan/app --for=condition=Ready --timeout=5m
kubectl get atlassecurityscans
NAME   READY   REASON    COMPLIANT   FINDINGS   HIGHEST   LAST SCAN   NEXT SCAN              AGE
app True Scanned False 3 HIGH 0s 2026-10-01T06:00:00Z 6s

COMPLIANT is False because a HIGH finding reached failOn. See Reading the Results.

kubectl get atlassecurityscans -o wide adds the trigger of the last scan, the schedule, and whether the resource is suspended.

Reading the Results​

Two conditions answer two different questions:

ConditionQuestionFalse means
ReadyDid the scan run?The last attempt failed; the reason says how. See Failures.
CompliantIs the database within policy, as of the last successful scan?A finding that is not waived reached failOn.

A finding is a result, not a failure: it moves Compliant and never Ready. Compliant is Unknown before the first scan, when failOn is unset, and once retries are exhausted. Use kubectl wait --for=condition=Ready to wait for a result, and kubectl wait --for=condition=Compliant as a deployment gate.

kubectl get atlassecurityscan app -o yaml
status:
conditions:
- lastTransitionTime: "2026-09-30T11:54:41Z"
message: 3 findings (1 HIGH, 2 ELEVATED) in 3 extensions
observedGeneration: 1
reason: Scanned
status: "True"
type: Ready
- lastTransitionTime: "2026-09-30T11:54:41Z"
message: 3 findings (1 HIGH, 2 ELEVATED) in 3 extensions
observedGeneration: 1
reason: Scanned
status: "False"
type: Reconciling
- lastTransitionTime: "2026-09-30T11:54:35Z"
message: ""
observedGeneration: 1
reason: Scanned
status: "False"
type: Stalled
- lastTransitionTime: "2026-09-30T11:54:41Z"
message: 1 finding at or above HIGH, highest HIGH; see atlassecurityreport/app
observedGeneration: 1
reason: PolicyViolated
status: "False"
type: Compliant
failed: 0
lastScan:
completionTime: "2026-09-30T11:54:41Z"
inputsHash: sha256:ed66fceb0749a62f75043390700525ae364d0b148671761f6a51c032bb30df5e
result: Succeeded
startTime: "2026-09-30T11:54:36Z"
trigger: Spec
triggeredBy: generation 1
lastSuccessfulTime: "2026-09-30T11:54:41Z"
nextScheduleTime: "2026-10-01T06:00:00Z"
observedGeneration: 1
reportRef:
name: app
summary:
driver: postgres
extensions: 3
highestLevel: HIGH
levels:
- count: 0
level: CRITICAL
- count: 1
level: HIGH
- count: 2
level: ELEVATED
- count: 0
level: NORMAL
total: 3
waived: 0

summary.levels always lists all four levels, highest first, so a count that drops to zero is reported as 0 rather than disappearing. The Reconciling and Stalled conditions follow the kstatus convention, so kstatus-aware tools such as Flux read the health of the resource without custom checks. Stalled is described in Failures.

The report​

The findings live in a separate AtlasSecurityReport with the same name, owned by the scan and replaced on every successful scan:

kubectl get atlassecurityreport app -o yaml
report:
completionTime: "2026-09-30T11:54:41Z"
extensions:
- hstore
- pg_trgm
- pgcrypto
policy:
failOn: HIGH
minSeverity: ELEVATED
serverVersion: "13.23"
startTime: "2026-09-30T11:54:36Z"
summary:
driver: postgres
extensions: 3
highestLevel: HIGH
levels:
- count: 0
level: CRITICAL
- count: 1
level: HIGH
- count: 2
level: ELEVATED
- count: 0
level: NORMAL
total: 3
waived: 0
trigger: Spec
vulnerabilities:
- cvssSeverity: MEDIUM
description: Buffer over-read in PostgreSQL pg_trgm index picksplit function reads
past end of a heap buffer. This might allow a table maintainer to infer limited
memory values, via the lossy signal of index split choices. Versions before
PostgreSQL 18.6, 17.11, 16.15, 15.19, and 14.24 are affected.
extension: pg_trgm
id: CVE-2026-14678
level: ELEVATED
suggestion: 'Upgrade the database engine to version 14.24 or later to resolve
CVE-2026-14678: the fix for extension "pg_trgm" ships in engine releases'
title: PostgreSQL pg_trgm picksplit reads past end of buffer
version: "1.5"
- cvssSeverity: MEDIUM
description: Cleartext storage in PostgreSQL pgcrypto disabled ciphers allows
a user to recover cleartext, via direct observation of the faulty ciphertext. The
OpenSSL version and OpenSSL configuration determine the disabled ciphers. If
the application accepts encrypted data as input, decryption will succeed even
with the wrong key. This in turn loses the modest protection from the Modification
Detection Code (MDC). Affected functions are pgp_sym_encrypt, pgp_sym_decrypt,
pgp_pub_encrypt, pgp_pub_decrypt, pgp_sym_encrypt_bytea, pgp_sym_decrypt_bytea,
pgp_pub_encrypt_bytea, and pgp_pub_decrypt_bytea. Versions before PostgreSQL
18.6, 17.11, 16.15, 15.19, and 14.24 are affected.
extension: pgcrypto
id: CVE-2026-14663
level: ELEVATED
suggestion: 'Upgrade the database engine to version 14.24 or later to resolve
CVE-2026-14663: the fix for extension "pgcrypto" ships in engine releases'
title: PostgreSQL pgcrypto, for OpenSSL-disabled ciphers, silently encrypts to
and decrypts from cleartext
version: "1.3"
- cvssSeverity: HIGH
description: Heap buffer overflow in PostgreSQL pgcrypto allows a ciphertext provider
to execute arbitrary code as the operating system user running the database. Versions
before PostgreSQL 18.2, 17.8, 16.12, 15.16, and 14.21 are affected.
extension: pgcrypto
id: CVE-2026-2005
level: HIGH
suggestion: 'Upgrade the database engine to version 14.21 or later to resolve
CVE-2026-2005: the fix for extension "pgcrypto" ships in engine releases'
title: PostgreSQL pgcrypto heap buffer overflow executes arbitrary code
version: "1.3"

The description of a finding is the text of its record, truncated to 1024 bytes. report.policy is the policy the report was graded with, and report.summary repeats the summary of the scan. A deleted report is recreated by the next successful scan, and deleting the scan deletes its report.

When Scans Run​

At least one of schedule and triggers must be set. A scan runs for any of the following causes, and status.lastScan.trigger records which one:

triggerCause
SpecThe resource was created, or its spec changed.
ManualThe db.atlasgo.io/scan-requested-at annotation was set to a new value.
ApplyA resource listed in triggers applied a new change.
ScheduleA schedule slot passed.
PolicyA waiver expired.

The table is in precedence order. When several causes are pending at once, one scan covers all of them and trigger names the one highest in the table.

On a schedule​

schedule is a five-field cron expression, or one of @hourly, @daily, @weekly, @monthly and @yearly. It is evaluated in timeZone, an IANA zone name that defaults to UTC, so a slot keeps its local time across daylight saving changes. The API server rejects @every and a TZ= or CRON_TZ= prefix: the first drifts rather than following the clock, and the zone belongs in timeZone.

Slots missed while the Operator was not running, or while scans kept failing, are caught up with a single scan, and a MissedSchedule event counts the slots it skipped before the one it covers. status.nextScheduleTime is the slot after the last one a successful scan covered, so it stays at a missed slot in the past until the catch-up succeeds. It is removed while the resource is suspended or stalled on an invalid spec.

After an apply​

triggers lists AtlasSchema and AtlasMigration resources in the same namespace, up to 32:

security-scan.yaml
spec:
schedule: "0 6 * * *"
timeZone: UTC
triggers:
- kind: AtlasSchema
name: app
- kind: AtlasMigration
name: app-migrations

A scan runs when a trigger applies a new revision: a new desired schema for an AtlasSchema, or a new last applied version for an AtlasMigration. A reconcile that applies nothing does not scan. A trigger that does not exist is reported with a TriggerNotFound event, and when the Operator is installed with labelSelector or watchNamespaces, a trigger must be managed by the same Operator instance.

On demand​

Set the db.atlasgo.io/scan-requested-at annotation to any new value. The value is echoed to status.lastHandledScanRequest once the scan succeeds, which gives kubectl wait something to wait on:

T=$(date -u +%FT%TZ)
kubectl annotate atlassecurityscan/app db.atlasgo.io/scan-requested-at="$T" --overwrite
kubectl wait atlassecurityscan/app --for=jsonpath="{.status.lastHandledScanRequest}=$T" --timeout=10m

Suspending​

spec.suspend: true stops scanning without deleting the resource or its report. Setting it back to false runs one scan that covers everything that became due in the meantime.

Policy​

spec.policy decides what is reported and what makes the database non-compliant:

FieldDescription
minSeverityThe lowest level reported: NORMAL, ELEVATED, HIGH or CRITICAL. NORMAL by default. Findings below it are not in the report.
failOnThe lowest level at which a finding that is not waived sets Compliant to False. When unset, findings are reported only.
ignoreWaivers for individual vulnerabilities, up to 256, each with an id, a reason, and an optional expirationTime.

The levels are those the Security Graph assigns, described in Severity Levels. The API server rejects a failOn below minSeverity, as a threshold under the reporting floor could never be reached.

Waivers​

A waiver keeps the finding in the report, marks it with the waiver's reason, and excludes it from the counts and the verdict. This differs from the CLI's --ignore flag, which drops the finding from the output. Waivers are listed under policy.ignore:

security-scan.yaml
spec:
policy:
minSeverity: ELEVATED
failOn: HIGH
ignore:
- id: CVE-2026-14678
reason: "pg_trgm is not reachable from the application role; SEC-1234"
expirationTime: "2026-12-31T00:00:00Z"
kubectl get atlassecurityscan app -o jsonpath='{.status.conditions[?(@.type=="Ready")].message}'
2 findings (1 HIGH, 1 ELEVATED) in 3 extensions; 1 waived

The waived finding in the report carries its waiver:

  - cvssSeverity: MEDIUM
description: Buffer over-read in PostgreSQL pg_trgm index picksplit function reads
past end of a heap buffer. This might allow a table maintainer to infer limited
memory values, via the lossy signal of index split choices. Versions before
PostgreSQL 18.6, 17.11, 16.15, 15.19, and 14.24 are affected.
extension: pg_trgm
id: CVE-2026-14678
level: ELEVATED
suggestion: 'Upgrade the database engine to version 14.24 or later to resolve
CVE-2026-14678: the fix for extension "pg_trgm" ships in engine releases'
title: PostgreSQL pg_trgm picksplit reads past end of buffer
version: "1.5"
waiver:
expirationTime: "2026-12-31T00:00:00Z"
reason: pg_trgm is not reachable from the application role; SEC-1234

When expirationTime passes, a scan runs with the Policy trigger and the finding counts again. status.activeWaivers lists the IDs of the waivers that were in effect when the last successful scan started. These are the only CVE identifiers the status holds.

Failures​

A scan that could not run sets Ready to False with one of the following reasons, and is retried with a linear backoff, 5 seconds longer per attempt. Compliant keeps the verdict of the last successful scan.

ReasonThe scan could not
ReadingInputsRead its Secrets, ConfigMaps or project configuration.
LoginFailedLog in to Atlas Cloud with the given token.
CLIErrorRun the Atlas CLI.
ScanFailedScan the database, for example because it was unreachable.
StoringReportWrite the AtlasSecurityReport.

The condition messages are fixed text, such as the database could not be scanned; attempt 3 of 20; see the operator log. The error of the CLI or the driver, which can contain the resolved address of a host, only goes to the Operator log.

After spec.backoffLimit retries fail (20 by default, so 21 attempts in total), the resource stalls: Stalled becomes True with the BackoffLimitExceeded reason, and Compliant becomes Unknown with ReportStale, as the last report may no longer describe the database. A stalled resource makes one attempt for each new schedule slot, apply, request, spec change or waiver expiry, instead of retrying on a timer. Set backoffLimit to 0 to retry without a limit.

Inputs that no retry can fix stall the resource without a backoff:

ReasonCauseTried again on
InvalidTimeZonetimeZone is not a known zone.A spec change.
InvalidScheduleschedule does not parse or never fires.A spec change.
InvalidTargetNo target is set, or a custom configuration is given without allowCustomConfig or envName.A spec change.
InvalidTargetThe configuration yields more than one database, or sets a rejected attribute.A new schedule slot, apply or request.

The first three are detected before any attempt. The last is detected during the attempt, and is tried again without a spec change because the configuration may be fixed outside the spec.

Access to Reports​

The scan holds only a summary. Its status never contains the connection URL, the extension names, a CVE identifier other than the waivers in spec.policy.ignore, or any CLI or driver error text, so read access to the scan does not reveal which extensions are vulnerable. The report lists every installed extension and each vulnerability found in it, so it is handled separately.

With rbac.aggregateClusterRoles: true, read access to AtlasSecurityScan is aggregated into the built-in view and edit roles. AtlasSecurityReport is not.

With rbac.create, the default, the chart ships a ClusterRole that grants read access to reports, named <release>-atlas-operator-securityreport-viewer, or <release>-securityreport-viewer when the release name already contains atlas-operator. It is aggregated into the built-in admin role by default:

values.yaml
rbac:
securityReports:
aggregateToAdmin: true
aggregateToEdit: false
aggregateToView: false

To let a group read the reports of one namespace without aggregating the role, bind it directly. For a release named atlas-operator:

report-readers.yaml
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: security-report-readers
namespace: app
subjects:
- kind: Group
name: security-team
apiGroup: rbac.authorization.k8s.io
roleRef:
kind: ClusterRole
name: atlas-operator-securityreport-viewer
apiGroup: rbac.authorization.k8s.io

Notifications and Custom Configuration​

A security block with notify endpoints can be given through a project configuration, which requires the Operator to be installed with allowCustomConfig=true. envName names the env of the configuration the scan runs with, and is required with a custom configuration: without it the resource stalls with InvalidTarget. The Operator sets the url of that env from the url, urlFrom or credentials of the scan, and a url the env sets itself takes precedence:

security-scan.yaml
spec:
envName: kubernetes
config: |
variable "slack_webhook" {
type = string
}
env "kubernetes" {
security {
notify {
http "slack" {
url = var.slack_webhook
body = jsonencode({ text = "${scan.count} findings (${scan.high} high)" })
}
}
}
}
vars:
- key: slack_webhook
valueFrom:
secretKeyRef:
name: slack
key: webhook

An endpoint that fails does not fail the scan. The findings are still graded and stored, and a ScanWarning event is recorded. The CLI sends the notification before the Operator applies spec.policy.ignore, so scan.count and the level counts include waived findings, and a scan whose findings are all waived still sends on the FINDINGS event.

Rejected attributes​

spec.policy is the only policy, and every extension of the database is scanned. The Operator rejects the following attributes on the env the scan runs with, using the InvalidTarget reason:

  • security.cve.min_severity, security.cve.ignore and exclude: they would drop findings or extensions before the Operator grades them.
  • security.min_severity and security.fail_on: they would define a second policy beside spec.policy.

Next Steps​