> ## Documentation Index
> Fetch the complete documentation index at: https://docs.rootly.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Private Agent Limits (Early Access)

> Reference for Private Agent limits across Kubernetes, observability, Kafka, Redis, Valkey, databases, internal MCP and HTTP providers, and AI context.

<Warning>
  **Early Preview:** Rootly Private Agent is under active development and available only to approved customers. Features, configuration, limits, and APIs may change before general availability. Confirm the approved agent and backend versions with your Rootly representative before production use.
</Warning>

Use this page as the canonical numeric reference for the early-access Private Agent. Setup guides link here rather than repeating policy defaults and ceilings. Confirm compatibility with your Rootly representative when upgrading either the agent or backend.

## Shared Runtime and AI Context

| Boundary                                          | Default or Maximum                  | Behavior                                                                                                                                                             |
| ------------------------------------------------- | ----------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `runtime.maximum_concurrency`                     | Default 8; configurable from 1 to 8 | Shared across Kubernetes, Prometheus, Loki, Tempo, Pyroscope, Argo CD, Kafka, Redis, Valkey, Elasticsearch, OpenSearch, PostgreSQL, MySQL, MCP, and HTTP invocations |
| Each registered input or output JSON schema       | 65,536 serialized bytes             | Oversized or non-serializable snapshots are rejected; the last accepted snapshot remains active                                                                      |
| Private-provider result entering AI model context | 128 KiB                             | Larger results are rejected before model context; narrow the request                                                                                                 |
| AI SRE Pod-log call                               | 64 KiB                              | Smaller than the local provider's log ceiling to preserve investigation context                                                                                      |
| Registered provider instances per combined agent  | 16                                  | Total across Kubernetes and every other adapter type                                                                                                                 |

Local provider limits and AI context limits are independent. A response that fits an adapter's local limit can still exceed the AI context limit. Request deadlines and cancellation also bound execution and result delivery.

## Kubernetes

| Local Boundary                    | Default | Hard Maximum |
| --------------------------------- | ------: | -----------: |
| Objects per list page             |     500 |        1,000 |
| Watch duration, seconds           |      30 |           60 |
| Encoded result size               |   2 MiB |        2 MiB |
| Pod-log bytes                     | 256 KiB |        1 MiB |
| Pod-log request duration, seconds |      30 |           60 |

The [Kubernetes guide](/private-agent-kubernetes#query-pod-logs-safely) explains targeted log arguments, truncation, permissions, and resource sizing. These bounds limit agent-side handling, not Kubernetes API server workload.

At most one Kubernetes entry may use in-cluster authentication. Additional
entries require mounted kubeconfigs. The shared 16-provider ceiling and runtime
concurrency limit apply across all Kubernetes clusters and other adapters. For
example, one Kubernetes cluster leaves room for 15 other provider instances,
while two Kubernetes clusters leave room for 14.

## Prometheus

Set these keys under an instance's `policy` mapping to tighten its limits.

| Local Policy                                          | Configuration Key           | Default | Hard Maximum |
| ----------------------------------------------------- | --------------------------- | ------: | -----------: |
| Concurrent invocations per instance                   | `maximum_concurrency`       |       2 |            8 |
| Response bytes after decompression                    | `maximum_result_bytes`      | 131,072 |      131,072 |
| Returned series or discovery entries                  | `maximum_series`            |     100 |        1,000 |
| Requested time window, seconds                        | `maximum_range_seconds`     |  86,400 |      604,800 |
| Range points per series                               | `maximum_points_per_series` |  11,000 |       11,000 |
| Request timeout, seconds, including admission waiting | `maximum_timeout_seconds`   |      15 |           30 |

Omitted or zero local policy values select defaults. Positive values may tighten the bounds. Explicit tool limits and timeouts must be positive and cannot exceed local policy. The shared runtime concurrency bound also applies.

| Input or Topology Boundary                                                |     Maximum |
| ------------------------------------------------------------------------- | ----------: |
| Provider instances per process, across all adapter types                  |          16 |
| PromQL expression or individual match selector                            | 8,192 bytes |
| Match selectors per request                                               |          20 |
| RFC3339 timestamp                                                         |    64 bytes |
| Metric or standard label name                                             |   255 bytes |
| `step_seconds`                                                            |      86,400 |
| Entire URL-encoded GET request target, including configured endpoint path | 8,192 bytes |

Metric-name and label-value discovery use GET-only upstream endpoints. Their aggregate request-target limit counts all selectors, percent-encoding, other parameters, and the configured endpoint path; individually valid arguments may need narrowing to fit. Metadata GETs share this bound. Query and label-name tools use form-encoded POST requests.

## Loki

Set these keys under an instance's `policy` mapping to tighten its limits.

| Local Policy                                          | Configuration Key         | Default | Hard Maximum |
| ----------------------------------------------------- | ------------------------- | ------: | -----------: |
| Concurrent invocations per instance                   | `maximum_concurrency`     |       2 |            8 |
| Response bytes after decompression                    | `maximum_result_bytes`    | 131,072 |      131,072 |
| Returned log, series, or discovery entries            | `maximum_entries`         |     100 |          100 |
| Requested time window, seconds                        | `maximum_range_seconds`   |   3,600 |       86,400 |
| Estimated bytes scanned per log or pattern query      | `maximum_scan_bytes`      |  10 GiB |      100 GiB |
| Pattern points                                        | `maximum_pattern_points`  |   1,000 |       11,000 |
| Request timeout, seconds, including admission waiting | `maximum_timeout_seconds` |      15 |           30 |

Omitted or zero values select defaults and never disable guardrails. Rootly AI SRE applies the lower of its reviewed ceiling and the instance's advertised local policy. At most 16 provider instances can be configured in one process across all adapter types.

| Input Boundary                                                        |     Maximum |
| --------------------------------------------------------------------- | ----------: |
| LogQL expression or stream selector                                   | 8,192 bytes |
| RFC3339 timestamp                                                     |    64 bytes |
| Standard label name                                                   |   255 bytes |
| `step_seconds`                                                        |      86,400 |
| Entire encoded request target, including the configured endpoint path | 8,192 bytes |

See the [Loki guide](/private-agent-loki#local-scope-and-scan-cost-controls) for selector enforcement, index-stat estimation limitations, sentinel pagination, and pattern-query behavior.

## Grafana Tempo

Set these keys under a Tempo instance's `policy` mapping to tighten its limits.

| Local Policy                                          | Configuration Key            | Default | Hard Maximum |
| ----------------------------------------------------- | ---------------------------- | ------: | -----------: |
| Concurrent invocations per instance                   | `maximum_concurrency`        |       2 |            8 |
| Response bytes after decompression                    | `maximum_result_bytes`       | 512 KiB |        2 MiB |
| Returned traces                                       | `maximum_traces`             |      20 |          100 |
| Returned spans per matching span set                  | `maximum_spans_per_span_set` |       3 |           10 |
| Attribute names or values                             | `maximum_attribute_values`   |     100 |        1,000 |
| Stale values inspected during attribute discovery     | `maximum_stale_values`       |   1,000 |       10,000 |
| Returned metric series                                | `maximum_series`             |     100 |        1,000 |
| Metric points per series                              | `maximum_points_per_series`  |   1,000 |       11,000 |
| Metric exemplars per series                           | `maximum_exemplars`          |      20 |          100 |
| Requested time window, seconds                        | `maximum_range_seconds`      |   3,600 |       86,400 |
| Request timeout, seconds, including admission waiting | `maximum_timeout_seconds`    |      15 |           30 |

Omitted or zero values select defaults. Rootly applies the lower of its reviewed ceiling and the instance's advertised local policy. Search, discovery, and metric windows must be at least one second. `tempo.get_trace` can omit its window or provide both timestamps together. At most 16 external provider instances can be configured in one process across all adapter types.

| Input Boundary                                                 |                                   Maximum |
| -------------------------------------------------------------- | ----------------------------------------: |
| Complete tool input                                            |                                    32 KiB |
| TraceQL query or filter                                        |                               8,192 bytes |
| TraceQL attribute                                              |                                 512 bytes |
| Hexadecimal trace ID                                           | 16 to 32 characters, even length, nonzero |
| RFC3339 timestamp                                              |                                  64 bytes |
| `step_seconds` for range metrics                               |                                    86,400 |
| Encoded request target, including the configured endpoint path |                               8,192 bytes |

See the [Grafana Tempo guide](/private-agent-tempo#query-and-result-boundaries) for exact-instance and tenant routing, TraceQL tools, security boundaries, health, and the Tempo 2.10/3.0 compatibility matrix.

## Pyroscope

Set these keys under a Pyroscope instance's `policy` mapping to tighten its limits.

| Local Policy                                          | Configuration Key           | Default | Hard Maximum |
| ----------------------------------------------------- | --------------------------- | ------: | -----------: |
| Concurrent invocations per instance                   | `maximum_concurrency`       |       2 |            8 |
| Response bytes after decompression                    | `maximum_result_bytes`      | 512 KiB |        2 MiB |
| Profile types                                         | `maximum_profile_types`     |     100 |        1,000 |
| Label values                                          | `maximum_label_values`      |     500 |        5,000 |
| Returned series                                       | `maximum_series`            |     100 |        1,000 |
| Points per series                                     | `maximum_points_per_series` |   1,000 |       10,000 |
| Flame-graph nodes                                     | `maximum_nodes`             |   1,024 |       10,000 |
| Requested time window, seconds                        | `maximum_range_seconds`     |   3,600 |       86,400 |
| Request timeout, seconds, including admission waiting | `maximum_timeout_seconds`   |      15 |           30 |

Selectors are limited to 8,192 bytes, profile type IDs to 1,024 bytes, labels
to 255 bytes, and label or grouping arrays to 16 unique entries. Up to 16
external provider instances can run in one process across provider types.
Profile queries require a selective positive label matcher and an explicit,
positive time window. `maximum_result_bytes` bounds the complete upstream JSON
response before normalization; the separate 128 KiB AI-context boundary still
applies. See the [Pyroscope guide](/private-agent-pyroscope#local-scope-and-result-controls)
for tenant routing, selector enforcement, tools, and compatibility.

## Argo CD

Set these keys under an Argo CD instance's `policy` mapping to tighten its limits.

| Local Policy                                          | Configuration Key         | Default | Hard Maximum |
| ----------------------------------------------------- | ------------------------- | ------: | -----------: |
| Concurrent invocations per instance                   | `maximum_concurrency`     |       2 |            8 |
| Encoded result bytes                                  | `maximum_result_bytes`    |   1 MiB |        2 MiB |
| Applications per response                             | `maximum_applications`    |     100 |          500 |
| Managed resources or tree nodes                       | `maximum_resources`       |     500 |        2,000 |
| Resource events                                       | `maximum_events`          |     200 |        1,000 |
| Destination clusters                                  | `maximum_clusters`        |     100 |          500 |
| Application history entries                           | `maximum_history_entries` |      50 |          200 |
| One desired or live manifest                          | `maximum_manifest_bytes`  |  64 KiB |      256 KiB |
| Request timeout, seconds, including admission waiting | `maximum_timeout_seconds` |      20 |           60 |

An instance can allow at most 100 distinct AppProjects. The complete upstream
response is capped at 16 MiB before parsing, then normalized into the lower
configured result budget. Application and resource names are limited to 512
bytes, filter expressions to 2 KiB, and the complete request target to 8 KiB.

Cluster inventory is a separate local opt-in. Secret and ConfigMap contents,
literal environment values, cluster credentials, managed fields, and
last-applied manifests are always removed. See the [Argo CD guide](/private-agent-argocd)
for upstream RBAC, tools, rotation, and compatibility.

## PostgreSQL and MySQL

Set these keys under a PostgreSQL or MySQL instance's `policy` mapping to tighten its limits.

| Local Policy                                          | Configuration Key         | Default | Hard Maximum |
| ----------------------------------------------------- | ------------------------- | ------: | -----------: |
| Concurrent invocations per instance                   | `maximum_concurrency`     |       2 |            8 |
| Returned rows                                         | `maximum_rows`            |     200 |        1,000 |
| Encoded tool-result bytes after driver decoding       | `maximum_result_bytes`    | 256 KiB |        1 MiB |
| SQL query bytes                                       | `maximum_query_bytes`     |  16 KiB |       64 KiB |
| Request timeout, seconds, including admission waiting | `maximum_timeout_seconds` |      15 |           30 |

Each instance supports at most 64 allowed schemas for typed catalog diagnostics. Arbitrary SQL is not parsed or scoped by this list; the database login's grants are its authorization boundary. These local limits apply in addition to those grants, the shared runtime concurrency ceiling, the invocation deadline, and the 128 KiB AI-context boundary.

`maximum_result_bytes` limits the encoded tool result returned to Rootly after the database driver has decoded each cell. It is an output/egress limit, not a pre-execution database or process-memory limit: the database and driver can still materialize one large value before the agent rejects the result. Limit the read-only identity to bounded investigative views, use database-side resource controls where appropriate, and size the agent container for the queries that identity can run.

See the [PostgreSQL and MySQL guide](/private-agent-databases#local-scope-and-limits) for database grants, TLS, SQL execution, and compatibility details.

## Kafka

Set these keys under a Kafka instance's `policy` mapping to tighten its limits.

| Local Policy                                          | Configuration Key           | Default | Hard Maximum |
| ----------------------------------------------------- | --------------------------- | ------: | -----------: |
| Concurrent invocations per instance                   | `maximum_concurrency`       |       2 |            8 |
| Encoded tool-result bytes                             | `maximum_result_bytes`      |   1 MiB |        2 MiB |
| Returned topics, groups, partitions, or other entries | `maximum_items`             |     500 |        5,000 |
| Returned records                                      | `maximum_messages`          |     100 |        1,000 |
| Aggregate returned record bytes                       | `maximum_message_bytes`     | 512 KiB |        2 MiB |
| Records inspected to satisfy a read                   | `maximum_scan_records`      |   5,000 |      100,000 |
| Request timeout, seconds, including admission waiting | `maximum_timeout_seconds`   |      20 |           60 |
| Transaction age classified as stuck, seconds          | `stuck_transaction_seconds` |     300 |       86,400 |

Each instance accepts 1 to 32 bootstrap servers and up to 100 allowed and 100
denied topic patterns. Patterns are limited to 255 bytes and Kafka-safe ASCII
glob characters. `maximum_scan_records` must be at least
`maximum_messages`. Omitted or zero numeric values select defaults and never
disable guardrails.

Message reads and message-value return are separate local opt-ins. Reads use
direct partition assignment, do not join a consumer group, and do not commit
offsets. All Kafka tools remain subject to the shared runtime concurrency and
128 KiB AI-context limits. See the [Kafka guide](/private-agent-kafka) for
authentication, topic scope, tool groups, and compatibility.

## Redis and Valkey

Set these keys under each Redis or Valkey instance's `policy` mapping to
tighten its limits. Both provider types use the same bounds.

| Local Policy                                          | Configuration Key         | Default | Hard Maximum |
| ----------------------------------------------------- | ------------------------- | ------: | -----------: |
| Concurrent invocations per instance                   | `maximum_concurrency`     |       2 |            8 |
| Encoded tool-result bytes                             | `maximum_result_bytes`    |  64 KiB |      256 KiB |
| Returned slowlog entries                              | `maximum_slowlog_entries` |      20 |          100 |
| Request timeout, seconds, including admission waiting | `maximum_timeout_seconds` |      10 |           30 |

The provider accepts at most 16 addresses per instance and at most one for
standalone mode. A tool input is at most 1 KiB; each individual `INFO` response
is capped at 64 KiB before field filtering. These limits do not bound work the
server performs to collect its own statistics, nor the raw arguments returned
by `SLOWLOG GET` before local redaction. A bounded diagnostic can also
be rejected by the shared 128 KiB AI-context limit. Rootly rejects returned
Redis/Valkey diagnostics that exceed the registered result-byte policy or
contain fields outside the reviewed output contract, including unredacted
slowlog arguments. See the [Redis and Valkey guide](/private-agent-redis-valkey)
for authentication, tools, and tested
versions.

## Elasticsearch and OpenSearch

Set these keys under an Elasticsearch or OpenSearch instance's `policy` mapping to tighten its limits.

| Local Policy                                          | Configuration Key         | Default | Hard Maximum |
| ----------------------------------------------------- | ------------------------- | ------: | -----------: |
| Concurrent invocations per instance                   | `maximum_concurrency`     |       2 |            8 |
| Response bytes after decompression                    | `maximum_result_bytes`    | 256 KiB |        1 MiB |
| Documents or native-query rows                        | `maximum_documents`       |     100 |          500 |
| Index or data-stream entries                          | `maximum_indices`         |     100 |          500 |
| Shard entries                                         | `maximum_shards`          |     200 |        1,000 |
| Requested Query DSL time window, seconds              | `maximum_range_seconds`   |   3,600 |      604,800 |
| Request timeout, seconds, including admission waiting | `maximum_timeout_seconds` |      15 |           30 |
| ES\|QL or PPL query bytes                             | `maximum_query_bytes`     |  32 KiB |       64 KiB |

Each instance requires 1 to 64 local index patterns after comma-separated entries are expanded. Each configured entry is limited to 1,024 bytes, and the combined entries are limited to 4 KiB so default health and discovery requests fit the request-target budget. Up to 128 simple source fields may be excluded, and search field selection is limited to 64 fields. Query DSL is also bounded by nesting, node, aggregation, bucket, and result limits enforced by the agent.

`maximum_indices` and `maximum_shards` are checked before every search-provider capability, including metadata discovery and health operations. OpenSearch Serverless enforces `maximum_indices` but does not expose a shard ceiling, so `maximum_shards` must be omitted for `aoss` providers. `maximum_result_bytes` applies to the tool result returned to Rootly. The agent requests only the fields needed for preflight index/shard scope validation and applies a separate fixed 1 MiB cap to that internal metadata, so tightening the visible result budget does not make a valid index scope unreadable. Query DSL `query` and `aggregations` objects are each bounded by `maximum_query_bytes`; the complete tool input has a separate hard 160 KiB allocation cap.

For Query DSL `search` and `count`, `maximum_timeout_seconds` is sent to the search engine and enforced as the agent's client deadline. The synchronous ES|QL and PPL APIs do not expose a general execution timeout, so their limit is a client-side request deadline; upstream work may take a short time to observe cancellation. Native queries are also constrained by their mandatory row cap and the query/result byte limits above.

Omitted or zero policy values select defaults and never disable guardrails. The Rootly backend applies the lower of its reviewed ceiling and the instance's advertised local policy. See the [Elasticsearch and OpenSearch guide](/private-agent-search#set-the-local-index-boundary) for the allowlist, field-exclusion, Query DSL, and native-language boundaries.

## MCP

Set these keys under an MCP instance's `policy` mapping to tighten its limits.

| Local Policy                                          | Configuration Key         | Default | Hard Maximum |
| ----------------------------------------------------- | ------------------------- | ------: | -----------: |
| Concurrent invocations per instance                   | `maximum_concurrency`     |       2 |            8 |
| Normalized result bytes                               | `maximum_result_bytes`    | 131,072 |        1 MiB |
| Request timeout, seconds, including admission waiting | `maximum_timeout_seconds` |      30 |           60 |
| Explicitly allowlisted tools                          | `maximum_tools`           |      64 |          128 |
| Normalized capability inventory per instance          | Not configurable          | 192 KiB |      192 KiB |

Each tool input schema is limited to 64 KiB, 16 levels, and 2,048 schema nodes. One server may advertise at most 512 tools during bounded discovery, but only the exact locally allowlisted tools can become capabilities. The shared limit of 16 provider instances applies across all adapter types.

Each MCP instance's normalized capability inventory is limited to 192 KiB. An oversized inventory makes that instance unhealthy and advertises no MCP tools, so it cannot prevent Kubernetes or another provider from registering. The complete provider snapshot is also checked locally against the 4 MiB control-plane limit with 64 KiB reserved for transport framing.

The direct non-SSE HTTP response is capped at 4 MiB while decoding. Every final normalized result must fit the lower configured result limit, and the existing 128 KiB AI-context limit still applies. See the [internal MCP guide](/private-agent-mcp#protocol-boundary) for schema reduction and unsupported client capabilities.

## Internal HTTP APIs

Set these keys under an HTTP instance's `policy` mapping to tighten its limits.

| Local Policy                                                      | Configuration Key         | Default | Hard Maximum |
| ----------------------------------------------------------------- | ------------------------- | ------: | -----------: |
| Concurrent requests per instance                                  | `maximum_concurrency`     |       2 |            8 |
| Serialized request body bytes                                     | `maximum_request_bytes`   |  32 KiB |      256 KiB |
| Serialized response result bytes, including metadata and encoding | `maximum_response_bytes`  |  64 KiB |       96 KiB |
| Request timeout, seconds, including admission waiting             | `maximum_timeout_seconds` |      15 |           30 |

Each provider permits 1 to 64 configured path prefixes, up to 2 KiB each, and up to 32 invocation-supplied request-header names. A query can contain up to 50 names and 100 scalar values per name. Query names are limited to 255 bytes, individual query and header values are limited to 8 KiB, and Rootly caps the encoded query at 24 KiB so the complete agent-side request target remains within 32 KiB after adding the base and invocation paths. Response headers are exposed only when locally allowlisted. The provider does not follow redirects and does not inherit environment proxy settings.

The health probe timeout is the lower of four seconds and the configured timeout, including any cold OAuth token exchange. Explicit HTTP health endpoints require a 2xx response; the default `/` probe also accepts 404 as origin reachability. Redirects, other non-2xx responses, transport failures, and TLS failures mark the provider unhealthy. See the [internal HTTP guide](/private-agent-http) for the fixed-origin boundary, credentials, and request behavior.

## Health Probes

External-provider health observations are cached for 15 seconds after success. Most instances have a bounded probe with a timeout of at most 2 seconds, independent of user-request admission, and reserved capacity. Elasticsearch and OpenSearch may request up to 4 seconds, capped by their lower configured request timeout, because AWS credential refresh and bounded index-access checks can require multiple calls. Their last successful index expression is tried first; otherwise at most eight access probes run concurrently. MCP failures can be retried without waiting for the healthy-cache interval. The runtime permits at most 16 concurrent provider probes so every configured provider instance fits in one probe wave within the 5-second readiness budget.

Invocation saturation alone does not mark an upstream unhealthy. A failed upstream probe can still affect readiness. See the [Prometheus guide](/private-agent-prometheus#health-and-compatibility), [Loki guide](/private-agent-loki#health-and-compatibility), [Grafana Tempo guide](/private-agent-tempo#health-and-compatibility), [Pyroscope guide](/private-agent-pyroscope#tls-health-and-compatibility), [Argo CD guide](/private-agent-argocd#health-compatibility-and-routing), [Kafka guide](/private-agent-kafka#health-routing-and-compatibility), [Redis and Valkey guide](/private-agent-redis-valkey#health-routing-and-compatibility), [Elasticsearch and OpenSearch guide](/private-agent-search#health-routing-and-compatibility), [PostgreSQL and MySQL guide](/private-agent-databases#health-and-compatibility), [internal MCP guide](/private-agent-mcp#health-discovery-and-changes), and [internal HTTP guide](/private-agent-http#health-and-troubleshooting) for status behavior and compatibility.

## Maintaining This Reference

Update this reference alongside changes to agent policy defaults, tool input validation, runtime health budgets, or Rootly's schema and AI-context guards. Keep setup examples illustrative, link to these tables for limits, and validate the examples against matching early-access builds. These limits are not a production load-capacity guarantee.
