kubectl.
Use this guide for live AI SRE investigation queries: Private Agent runs bounded
queries against the cluster selected for a responder-started or automatically
triggered investigation. This is not continuous collection. For automatic
cluster-event ingestion as pulses, use the existing
Kubernetes integration and its
configuration guide. That workflow uses kubewatch
and webhooks; its setup and credentials are separate from Private Agent.
For responder-started investigations, Rootly applies the initiating user’s Private Agent permissions; sensitive reads such as Pod logs require an owner or admin. An automatic AI SRE investigation has no initiating user. When both AI SRE and Private Agent are enabled, the ai-sre system actor can use registered Kubernetes capabilities under the agent’s local policy and the selected provider’s Kubernetes RBAC. Scope both for unattended access.
This makes the same agent compatible with conformant clusters including EKS, GKE, AKS, self-managed Kubernetes, and local Kind clusters. Kubernetes API discovery selects a version served by the connected cluster; callers can optionally request a specific API group and version.
The automated Kind test matrix covers Kubernetes 1.34, 1.35, and 1.36. Discovery
does not make every Kubernetes version or optional API supported: a resource
must be served by the cluster, supported by the provider, and permitted by local
policy and Kubernetes RBAC. Argo Rollouts, KEDA, Istio, Gateway API, and resource
metrics require their corresponding APIs to be installed.
Supported read operations
The provider supports get, list, and bounded watch operations for:- Pods, Deployments, ReplicaSets, StatefulSets, DaemonSets
- Jobs and CronJobs
- Events, Nodes, Namespaces
- Services, Endpoints, and EndpointSlices
- ConfigMaps
- PersistentVolumes and PersistentVolumeClaims
- ResourceQuotas and LimitRanges
- StorageClasses and CustomResourceDefinitions
- HorizontalPodAutoscalers and PodDisruptionBudgets
- Ingresses, IngressClasses, NetworkPolicies, IPAddresses, ServiceCIDRs, and supported Gateway API resources
- Pod and Node resource metrics
- APIServices, admission policies and bindings, and admission webhook configurations
- PriorityClasses and RuntimeClasses
- CSIDrivers, CSINodes, CSIStorageCapacities, VolumeAttachments, and VolumeAttributesClasses
- DeviceClasses, ResourceClaims, ResourceClaimTemplates, and ResourceSlices
- Argo Rollouts and KEDA ScaledObjects
- Resources discovered in the
networking.istio.ioAPI group
kubernetes.api_resources to
discover the served versions and supported get, list, and watch verbs.
It accepts optional group, resource, namespaced, and limit filters;
group: "" explicitly selects the core API. Results are bounded to 500 entries
and indicate partial discovery or truncation. For metrics, explicitly select
group: "metrics.k8s.io" and resource: "pods" or resource: "nodes"; those
resource names also exist in the core API and otherwise can resolve to ordinary
objects rather than metrics.
Pod logs are a separate sensitive_read capability. ConfigMap data and binaryData are removed from generic results. Literal container environment-variable values are redacted from Pods and supported workload templates; variable names and valueFrom references remain visible. The kubectl.kubernetes.io/last-applied-configuration annotation is removed from returned objects because it can contain an older, unredacted manifest. Kubernetes Secrets and mutation capabilities are not supported. Redaction is not a guarantee that all customer-authored metadata or log content is non-sensitive; scope access accordingly.
Diagnose authorization
Two safe-read capabilities inspect the Kubernetes identity configured for the selected provider:kubernetes.self_access_reviewchecks a supported resource operation.getrequires a nonemptyname;listandwatchmust omitname. Provide a namespace for namespaced resources. Customer-local policy and served-verb checks apply before the request reaches Kubernetes. The only supported subresource ispods/logwithget, and local Pod-log access must be enabled.kubernetes.self_rules_reviewrequires an allowed namespace and returns raw effective Kubernetes RBAC, includingincompleteandevaluation_errorindicators. Its output can include write verbs, unsupported resources, and cluster-scoped rules. These are not additional operations the agent can execute, and an incomplete response must not be treated as a complete access inventory.
Configure local policy
Mount a non-secret YAML configuration file and setROOTLY_PRIVATE_AGENT_CONFIG_FILE to its path. Secret values do not belong in this file.
Multi-cluster support intentionally changes the pre-GA
providers.kubernetes
setting from a mapping to an array. There is no compatibility shim: deploy a
matching Private Agent and Helm chart release, and convert even a
single-cluster configuration to the array form shown above.providers.kubernetes or set it to an empty list. With the Helm chart, set
providers.kubernetes: [] and rbac.create: false so the chart also omits
Kubernetes RBAC and ServiceAccount token automounting.
Choose a stable provider id that is unique within the Rootly account, such as the cluster name plus region. Rootly rejects the same ID from a second active Private Agent so work cannot be routed ambiguously between clusters. Revoke a retired installation before reusing its provider ID.
An empty allowed_namespaces list denies all namespaced operations. Use "*" only when you intentionally want every namespace. Cluster-scoped access and Pod logs are disabled independently by default.
Connect multiple clusters
Each list entry registers an independently routed cluster. An entry withoutkubeconfig_file uses the agent Pod’s ServiceAccount; at most one entry may use
that mode. Add other clusters with an absolute mounted kubeconfig path and an
optional exact context. When context is omitted, the agent uses the file’s
current-context; it never guesses among the other contexts in that file:
KUBECONFIG, a home-directory config, or other ambient
contexts. Rootly can route to an advertised provider ID but cannot supply or
override a cluster endpoint, path, context, credential, or local policy.
EKS, GKE, AKS, self-managed Kubernetes, and Kind expose the same Kubernetes API
surface used by the provider. Authentication is separate: the published
distroless image does not include aws, gcloud, kubelogin, or arbitrary
kubeconfig exec helpers. The agent rejects exec and legacy auth-provider
authentication stanzas instead of invoking environment-dependent plugins.
The selected API server must use HTTPS with certificate verification enabled;
plaintext endpoints, embedded URL credentials, query strings, fragments, and
insecure-skip-tls-verify are rejected.
Prefer one in-cluster agent per cluster for the smallest credential boundary.
For a centralized deployment, mount a customer-managed kubeconfig with static
credentials or mounted token, CA, and client-certificate files, and allow
network access from the agent to every configured API endpoint. Relative file
references resolve from the kubeconfig’s directory.
Policies and health are reported per cluster. If one API becomes unavailable,
its provider becomes unhealthy without changing another cluster’s routing or
health. Overall /readyz remains conservative and reports not ready while any
configured provider is not healthy.
Configure Kubernetes RBAC
Local policy is not a substitute for Kubernetes RBAC. For an in-cluster provider, the Kubernetes identity is the agent Pod’s ServiceAccount. For a kubeconfig-backed provider, it is the user or service identity authenticated by that kubeconfig. Grant each configured identity only the resources its investigations need in that target cluster:- Use a namespace
RoleandRoleBindingwhen access is limited to one namespace. - For several namespaces, create a
RoleandRoleBindingin each allowed namespace, or define one read-onlyClusterRolefor namespaced resources and reference it from a separateRoleBindingin each allowed namespace. A namespacedRolecannot be reused from a different namespace. For an in-cluster provider, each binding’s ServiceAccount subject must reference the namespace where the agent runs. For a kubeconfig-backed provider, bind the user, group, or service identity represented by its credential. - For cluster-scoped resources such as Nodes, Namespaces, PersistentVolumes, StorageClasses, or CRD definitions, use a
ClusterRolewith aClusterRoleBindingonly when those reads are required and permitted by local policy. NamespaceRoleBindings do not grant cluster-scoped access. - Grant
getonpods/logonly when Pod-log access is enabled. - For authorization diagnostics, ensure the configured Kubernetes identity can
createselfsubjectaccessreviewsandselfsubjectrulesreviewsinauthorization.k8s.io. These cluster-scoped API calls evaluate access; they do not persist objects or grant permission to modify workloads or RBAC. - Both review resources use cluster-level API endpoints. For
SelfSubjectRulesReview, the requiredspec.namespaceselects the namespace whose rules are evaluated; it does not make the review resource namespaced. Verify existing self-review permissions before adding bindings. The Kubernetes client API reference lists these create endpoints without a/namespaces/{namespace}path segment. - Do not grant Secrets, workload write verbs, RBAC administration, exec, attach, port-forward, impersonation, or service-account token creation.
get, list, and watch. The provider cannot exceed either RBAC or its local YAML policy.
Mount enrollment and credential state
The enrollment token is short-lived and used once. Mount it from a Secret at the configured token path. The agent stores its rotating credential state on durable writable storage so Pod restarts do not require a new enrollment token. The production image runs as UID and GID65532. When a Secret volume uses mode 0440, set a Pod-level group so the non-root process can read it. This is a Pod-spec fragment, not a complete installation manifest; the Secret, configuration mount, and writable persistent credential volume must also be supplied:
fsGroup or equivalent group ownership, a 0440 Secret file owned by root with group root is not readable by this non-root image. Keep the fsGroup: 65532 setting with the mount shown above; do not make the token world-readable to work around the issue.
Query Pod logs safely
Enable Pod logs only when investigations require them. Each request supports KubernetesPodLogOptions, including container, previous, since_seconds, since_time, timestamps, tail_lines, limit_bytes, and stream.
Kubernetes does not provide an end-time argument for Pod logs. Keep requests targeted with:
since_secondsor RFC3339since_timetail_lineslimit_bytes- A request timeout
Resource sizing and performance
Start a combined agent with requests around100m CPU and 128Mi memory and limits around 1 CPU and 512Mi memory, then tune from observed usage. These are starting points, not universal requirements.
Load is controlled by independent runtime concurrency, list pagination, watch duration, encoded result size, log byte/time, and AI-context bounds. See Private Agent Limits for defaults and hard ceilings. Request deadlines and cancellation propagate into Kubernetes calls.
Large object metadata, high event volume, and concurrent watches can still increase CPU and memory. Lower runtime.maximum_concurrency for a small Pod or an API server with a tight request budget; excess work stays queued in Rootly until capacity is available. This reduces peak CPU, memory, and Kubernetes API pressure at the cost of parallel investigation throughput. Monitor container CPU, working-set memory, restarts, readiness, Kubernetes API throttling, invocation latency, and truncated-result counts. Narrow selectors, namespaces, time windows, and page sizes before increasing resources.
Network requirements
Allow outbound TCP 443 from the agent to its configured Rootly gRPC origin and access from the Pod to every configured cluster’s Kubernetes API endpoint. The default origin ishttps://connect.rootly.com. The control-plane channel requires gRPC over TLS and HTTP/2, including long-lived bidirectional streams; opening TCP 443 alone is not sufficient if a firewall or proxy downgrades or blocks HTTP/2. Enrollment, token refresh, provider registration, and work use this same endpoint. No inbound internet access, public Service, or Ingress is required.
If Rootly supplies a regional or custom endpoint, set rootly.grpc_url (or ROOTLY_PRIVATE_AGENT_GRPC_URL) and allow outbound access to that origin instead of connect.rootly.com. The value must be an origin without a path, query, or credentials. When omitted, legacy rootly.api_url selects the same gRPC endpoint, defaulting to https://connect.rootly.com. This override changes only the Private Agent control-plane transport origin; it does not change the region where Rootly AI processes or retains investigation data. See Data Privacy for Rootly AI for the current processing-region policy.
For enterprise proxies or TLS inspection, validate HTTP CONNECT and streaming behavior with your Rootly representative. Use a configured rootly.ca_bundle_file for an additional trusted CA when needed; certificate and hostname verification remain enabled. There is no HTTP polling fallback if gRPC is blocked.
TLS certificates are validated normally. Plain HTTP is rejected unless an explicit development-only override is enabled.