Skip to content

Events & Troubleshooting

Access Kubernetes events with descriptions and troubleshooting guidance. There is currently no container log-streaming feature — see the Pod Logs note below.

What Are Events?

Kubernetes generates events as things happen in your cluster. Each event includes context and guidance:

  • What happened: Description of the event
  • Why it matters: Impact on your applications
  • What to do: Recommended actions if needed
  • Context: Related resources and timeline

Event Context Provided

Events include descriptions and troubleshooting steps to help you quickly identify and resolve issues.

Event Categories

✅ Normal Events (Good News)

These events mean things are working as expected:

What you'll see:

  • "Pod successfully started on node worker-1"
  • "Container image pulled and ready to run"
  • "Application deployment completed successfully"
  • "Storage volume attached and ready"

Why we show these: To confirm your operations completed successfully and provide an audit trail.

⚠️ Warning Events (Pay Attention)

These events indicate potential issues that haven't caused failures yet:

What you'll see (with explanations):

  • "Pod waiting to be scheduled"

    • What it means: No node has enough resources for this pod
    • What to do: Check if you need to add nodes or reduce resource requests
  • "Image pull is slow"

    • What it means: Container image is taking a while to download
    • What to do: Usually resolves itself. If persistent, check network connectivity
  • "Health check failing"

    • What it means: Your application isn't responding to readiness probes
    • What to do: Check application logs for startup issues

Only Two Real Event Types

Kubernetes (and our API) only classifies events as Normal or Warning — there is no separate "Error" type. The scenarios below surface as Warning events; we still call out the more serious ones separately here because they usually need immediate action.

🔴 Serious Warning Events (Action Required)

These Warning events indicate active problems:

What you'll see (with troubleshooting):

  • "Container crashed with exit code 1"

    • What it means: Your application exited with an error
    • What to do: Check your application's own logging/monitoring for the error message (this platform doesn't provide container log access)
  • "Out of memory (OOMKilled)"

    • What it means: Container used more memory than its limit
    • What to do: Increase memory limits or optimize your application
  • "Cannot pull image"

    • What it means: Can't download the container image
    • What to do: Verify image name, registry access, and credentials

Event Sources

Events are generated by various Kubernetes components:

Scheduler

  • Pod scheduling decisions
  • Node selection
  • Resource constraints

Kubelet

  • Container lifecycle
  • Volume operations
  • Node status changes

Controller Manager

  • Deployment rollouts
  • ReplicaSet scaling
  • Job execution

API Server

  • Resource creation/deletion
  • Authentication/Authorization events

Event Information

Each event contains:

  • Type: Normal or Warning
  • Reason: Event classification (e.g., Failed, Started, Killing)
  • Message: Detailed description
  • Source: Component that generated the event
  • Object: Related Kubernetes resource
  • Timestamp: When the event occurred
  • Count: Number of occurrences

Event Monitoring

Resource Events

Monitor events for specific resources to track their lifecycle and identify issues.

Configuration:

yaml
clusterPirate:
  monitoring:
    resourceEventsEnabled: true

System Events

Track cluster-wide and node-level events to monitor infrastructure health.

Configuration:

yaml
clusterPirate:
  monitoring:
    systemEventsEnabled: true

Common Event Scenarios

Pod Failures

Failed Scheduling

  • Reason: FailedScheduling
  • Common Causes: Insufficient resources, node selectors, taints/tolerations
  • Resolution: Check node capacity, adjust resource requests, verify node labels

Image Pull Errors

  • Reason: Failed, ErrImagePull, ImagePullBackOff
  • Common Causes: Invalid image name, missing credentials, network issues
  • Resolution: Verify image name, check image pull secrets, test registry connectivity

Container Crashes

  • Reason: CrashLoopBackOff, Error
  • Common Causes: Application errors, missing dependencies, configuration issues
  • Resolution: Check container logs, verify environment variables, review application code

OOM Kills

  • Reason: OOMKilled
  • Common Causes: Insufficient memory limits, memory leaks
  • Resolution: Increase memory limits, profile application memory usage, fix memory leaks

Volume Issues

Mount Failures

  • Reason: FailedMount
  • Common Causes: Volume not available, incorrect PVC configuration, storage class issues
  • Resolution: Verify PVC exists, check storage class, ensure volume is bound

Volume Full

  • Reason: VolumeResizeFailed
  • Common Causes: Disk space exhausted, volume resize not supported
  • Resolution: Clean up disk space, resize volume, migrate to larger volume

Readiness/Liveness Failures

Probe Failures

  • Reason: Unhealthy
  • Common Causes: Application not ready, incorrect probe configuration, network issues
  • Resolution: Check application startup time, adjust probe settings, verify endpoint availability

Pod Logs

Not Implemented

There is no container log-streaming feature anywhere in the platform today — no way to view stdout/stderr, historical logs, or filter log content. Everything below is a planned capability, not something you can use now. Don't confuse this with the unrelated, real deployed-application logs endpoint in the Managed Application Platform, which is a different feature for a different resource type.

Accessing Logs (planned)

Logs are available for all containers in running and recently terminated pods.

Via Web Console:

  1. Navigate to cluster in portal
  2. Select namespace and pod
  3. Choose container (if multiple)
  4. View real-time logs

Log Features (planned)

Real-time Streaming

  • Live tail of container stdout/stderr
  • Automatic updates as new logs are written

Historical Logs

  • Access logs from previous container runs
  • View logs from terminated containers

Filtering

  • Search log content
  • Filter by timestamp
  • Filter by log level (if structured)

Log Retention (planned)

  • Active Containers: Logs available while container is running
  • Terminated Containers: Logs retained based on Kubernetes configuration
  • Pod Deletion: Logs are lost when pod is deleted

Use Cases

Troubleshooting Application Issues

  1. Check Pod Events: Identify scheduling or startup issues
  2. Review Container Logs: Look for application errors or exceptions (via your own logging setup — not available through this platform)
  3. Monitor Resource Events: Track deployment updates and rollouts
  4. Examine System Events: Identify infrastructure problems

Debugging Crashes

  1. Find OOM Events: Check for memory-related kills
  2. Review Exit Codes: Understand how containers terminated
  3. Analyze Error Patterns: Identify recurring issues

Monitoring Deployments

  1. Track Rollout Events: Monitor deployment progress
  2. Identify Pod Failures: Catch issues during updates
  3. Verify Configuration: Ensure correct settings applied
  4. Watch Resource Updates: Track replica changes

Audit Trail

  1. Resource Creation: Track when resources were created
  2. Configuration Changes: Monitor updates to workloads
  3. Access Events: Review authentication/authorization events
  4. Deletion Events: Track resource cleanup

Kubernetes Events

Events come from a dedicated, cluster-wide endpoint — fetching a specific resource (see Kubernetes Resources) never embeds events in its response.

List Cluster Events

http
GET /v1/workspaces/{workspaceId}/observability/{clusterId}/kubernetes-events
Authorization: Bearer <access-token>

This is a paginated list (see Pagination) — the response body is a bare array:

json
[
  {
    "workspaceId": "...",
    "clusterId": "...",
    "resourceId": "...",
    "eventId": "...",
    "firstTimestamp": "2024-01-15T10:30:00Z",
    "lastTimestamp": "2024-01-15T10:30:00Z",
    "count": 1,
    "message": "Started container nginx",
    "reason": "Started",
    "type": "NORMAL",
    "involvedObject": {
      "kind": "Pod",
      "name": "nginx-abc123",
      "uid": "...",
      "namespace": "production"
    },
    "source": {
      "component": "kubelet",
      "host": "worker-1"
    }
  }
]

type is uppercase-only (NORMAL or WARNING).

Get a Single Resource (No Events)

http
GET /v1/workspaces/{workspaceId}/observability/{clusterId}/kubernetes-proxy/namespaces/{namespace}/pod/{podName}
Authorization: Bearer <access-token>

Response:

json
{
  "clusterId": "...",
  "workspaceId": "...",
  "resource": {
    /* raw pod manifest, managedFields stripped */
  }
}

Best Practices

Event Monitoring

  • Enable both resource and system events for complete visibility
  • Regularly review warning events to catch issues early (there is no alerting engine yet — see Alert Reference)

Log Management

  • Implement structured logging in applications
  • Include correlation IDs for request tracing
  • Use appropriate log levels (DEBUG, INFO, WARN, ERROR)
  • Avoid logging sensitive information

Troubleshooting Workflow

  1. Start with events to identify the problem type
  2. Review your application's own logs for application-specific details (not available through this platform)
  3. Check resource configuration for misconfigurations
  4. Examine metrics for resource constraints
  5. Review cluster-wide events for infrastructure issues