Saturday, July 25, 2026

OOMKilled in Kubernetes: Understanding Exit Code 137 and How to Fix It

If you've worked with Kubernetes for some time, you've probably come across a pod that suddenly restarts with the reason OOMKilled and an exit code of 137. It usually happens without much warning. One moment the application is running normally, and the next moment Kubernetes terminates the container and starts a new one.

The first reaction is often to increase the memory limit and redeploy the application. Sometimes that works, but in many cases the same issue returns after a few hours or days because the actual problem was never investigated.

In most production environments, OOMKilled isn't the real problem. It's an indication that the application consumed more memory than it was allowed to use. Our objective shouldn't be to simply allocate more memory. Instead, we should understand why the application exceeded its limit in the first place.

In this article, we will understand what OOMKilled means, why Kubernetes reports Exit Code 137, how to investigate memory-related issues and the best practices to prevent them from happening again.

What Does OOMKilled Mean?

OOM stands for Out Of Memory.

Every container running in Kubernetes has resource limits that define the maximum amount of memory it can consume. If the application crosses that limit, the Linux kernel immediately terminates the process to protect the node from running out of memory.

Kubernetes detects that the container has stopped unexpectedly and restarts it according to the pod's restart policy.

If the application continues exceeding the memory limit after every restart, the pod may eventually enter a CrashLoopBackOff state.

Understanding this behaviour is important because Kubernetes isn't killing the application. The operating system is. Kubernetes simply reports what happened and starts a new container.

Understanding Exit Code 137

When a Linux process exits, it returns an exit code.

For an OOMKilled container, Kubernetes usually reports something similar to the following:

Last State:
  Terminated

Reason:
  OOMKilled

Exit Code:
  137

Exit Code 137 indicates that the process was terminated using the SIGKILL signal.

In Kubernetes, this almost always means one of the following:

  • The application exceeded its configured memory limit.

  • The Linux Out Of Memory Killer terminated the process.

  • Kubernetes restarted the container after it exited.

Whenever we see Exit Code 137, memory usage should be the first thing we investigate.

Confirm That the Pod Was OOMKilled

The easiest way to verify the reason is by describing the pod.

kubectl describe pod <pod-name>

Look for the Last State section.

Last State:
  Terminated

Reason:
  OOMKilled

Exit Code:
  137

If both the reason and exit code match the output above, we've confirmed that memory exhaustion caused the container to terminate.

Before making any configuration changes, we should understand why it happened.

Review the Resource Configuration

The next step is checking the memory requests and limits configured for the container.

Run:

kubectl get pod <pod-name> -o yaml

Locate the resources section.

resources:
  requests:
    memory: "512Mi"
    cpu: "250m"

  limits:
    memory: "1Gi"
    cpu: "500m"

There is often confusion between requests and limits.

A memory request determines the amount of memory Kubernetes reserves for the container during scheduling.

A memory limit defines the maximum memory the application is allowed to consume.

Once the application crosses that limit, the Linux kernel terminates the process immediately.

Setting memory limits too low is one of the most common reasons for OOMKilled pods.

Check Current Memory Usage

Before increasing memory limits, we should understand how much memory the application is actually consuming.

If Metrics Server is installed, run:

kubectl top pod

or

kubectl top pod <pod-name>

Example:

NAME              CPU(cores)   MEMORY(bytes)
payment-api       210m         985Mi

If the container has a memory limit of 1Gi and is already consuming around 985Mi, even a small increase in workload may cause the application to exceed its limit.

Monitoring memory usage gives us a much clearer picture than simply guessing.

Check the Previous Logs

When Kubernetes restarts a container, the current logs may only contain startup information.

The actual problem usually exists in the previous container.

Run:

kubectl logs <pod-name> --previous

Depending on the application, we may find messages related to:

  • Memory allocation failures

  • Large file uploads

  • Garbage collection warnings

  • Cache growth

  • Unexpected spikes in workload

Checking the previous logs has helped us identify the root cause of many production incidents.

Common Reasons for OOMKilled

Although every application behaves differently, most OOMKilled incidents fall into a few common categories.

Memory Leaks

Applications that continuously allocate memory without releasing it eventually consume all available memory.

Initially everything appears normal.

After running for several hours or days, memory usage gradually increases until the container reaches its configured limit.

Monitoring memory growth over time usually helps identify this pattern.

Processing Large Files

Applications processing PDFs, images, videos or large Excel files often require much more memory than expected.

If the entire file is loaded into memory, temporary spikes can exceed the configured limit.

Whenever possible, processing files in smaller chunks significantly reduces memory consumption.

Loading Large Datasets

Another common mistake is loading an entire dataset into memory before processing it.

Instead of loading thousands of records at once, processing them in batches or using streaming techniques keeps memory usage under control.

Traffic Spikes

Applications that perform well under normal traffic may consume considerably more memory during peak usage.

Higher request volumes often result in additional objects being created, more active database connections and larger caches.

Without sufficient memory planning, the container eventually exceeds its limit.

Incorrect Resource Configuration

Sometimes the application itself isn't the problem.

The configured memory limit simply doesn't reflect the application's actual workload.

Development and testing environments often use much smaller datasets than production, making it difficult to estimate realistic memory requirements.

Should We Simply Increase the Memory Limit?

Increasing the memory limit may stop the restarts temporarily, but it shouldn't be the first solution.

Before changing resource limits, it's worth asking a few questions.

  • Did the issue start after a recent deployment?

  • Has application traffic increased?

  • Are larger files being processed?

  • Has a new feature introduced additional memory usage?

  • Is memory usage continuously increasing over time?

Answering these questions often leads us to the actual root cause instead of masking the problem.

Review Recent Deployments

If the application was working correctly yesterday but started failing after a deployment, compare the recent changes.

Run:

kubectl rollout history deployment <deployment-name>

Recent code changes may have introduced:

  • Larger in-memory caches

  • Additional background workers

  • New libraries

  • Bigger response payloads

  • Changes in data processing

If production is impacted, rolling back to the previous deployment can restore service while the investigation continues.

Best Practices to Prevent OOMKilled

Preventing memory issues is much easier than troubleshooting them during an outage.

Some practices that have worked well across production environments include:

  • Configure realistic memory requests and limits.

  • Continuously monitor CPU and memory usage.

  • Process large files in batches.

  • Stream large datasets whenever possible.

  • Optimise application caching.

  • Load test applications before production releases.

  • Review memory consumption after every major deployment.

  • Configure Horizontal Pod Autoscaler where appropriate.

Small improvements in memory management often make a significant difference to application stability.

Common Mistakes

These are some of the mistakes we see most frequently while investigating OOMKilled incidents.

  • Increasing memory limits without identifying the root cause.

  • Ignoring historical memory usage.

  • Deploying applications without load testing.

  • Assuming Kubernetes is responsible for application crashes.

  • Forgetting to review previous container logs.

  • Using the same resource configuration for every environment.

Avoiding these mistakes usually makes troubleshooting much faster.

OOMKilled is one of the most common Kubernetes issues, but it's also one of the easiest to diagnose once we understand what Exit Code 137 represents.

Rather than treating OOMKilled as the actual problem, we should treat it as an indication that the application consumed more memory than its configured limit.

A structured troubleshooting process always produces better results than making configuration changes based on assumptions. Confirm the reason using kubectl describe, review the configured resource limits, analyse memory usage, inspect the previous logs and compare recent deployments before increasing memory.

In many cases, Kubernetes has already provided everything we need to identify the root cause. We simply need to collect the information in the right order and let the evidence guide our investigation.

Why Your Kubernetes Pod Keeps Restarting: A Complete Debugging Guide

If you've been working with Kubernetes for any length of time, you've probably encountered a pod that refuses to stay running. One moment it's starting successfully, and the next moment it's restarting. After a few attempts, Kubernetes reports a CrashLoopBackOff status, leaving many engineers wondering what went wrong.

The first reaction is often to delete the pod and hope Kubernetes creates a healthy replacement. While this may appear to solve the problem temporarily, it rarely fixes the actual issue. More importantly, deleting the pod too quickly can remove valuable information that would have helped identify the root cause.

The good news is that Kubernetes almost always tells us why a pod is restarting. We simply need to know where to look and follow a structured troubleshooting approach instead of making assumptions.

In this article, we will walk through the exact process we can use to diagnose and resolve Kubernetes pods that keep restarting.

Step 1: Check the Pod Status

Whenever a pod starts restarting, our first objective is to understand what Kubernetes already knows about it.

The first command we should run is:

kubectl get pods

Example output:

NAME                    READY   STATUS             RESTARTS   AGE
payment-api-65d8fd      0/1     CrashLoopBackOff   12         18m

Pay close attention to these three columns:

  • STATUS

  • RESTARTS

  • AGE

A restart count of 10 or 20 immediately tells us the application has been crashing repeatedly and Kubernetes has been trying to recover it.

Some common pod statuses include:

StatusMeaning
RunningApplication is healthy
PendingWaiting to be scheduled
CrashLoopBackOffContainer keeps crashing after startup
ImagePullBackOffUnable to pull the container image
ErrImagePullImage download failed
CompletedJob completed successfully
TerminatingPod is shutting down

Once we've confirmed the pod is restarting, the next step is to gather more information.

Step 2: Describe the Pod

One of the most useful commands in Kubernetes troubleshooting is:

kubectl describe pod <pod-name>

Many engineers immediately jump to application logs, but kubectl describe provides far more context.

It includes information such as:

  • Current container state

  • Previous container state

  • Restart count

  • Mounted volumes

  • Environment variables

  • Probe failures

  • Scheduling information

  • Events

The Events section at the bottom deserves special attention because Kubernetes records everything significant that happened to the pod.

For example:

Warning  Unhealthy
Readiness probe failed

Warning  BackOff
Back-off restarting failed container

In many situations, the Events section already points us in the right direction before we even inspect the application logs.

Step 3: Read the Application Logs

Once we've reviewed the pod details, it's time to inspect the application logs.

kubectl logs <pod-name>

If the pod has already restarted, don't forget to check the logs from the previous container instance:

kubectl logs <pod-name> --previous

This is one of the most overlooked commands in Kubernetes.

When a container crashes and restarts, the current logs may only show startup messages. The actual exception often exists only in the previous container's logs. We've solved many production issues simply by checking kubectl logs --previous.

Now that we've gathered the basic information, we can start identifying the actual reason behind the restarts.

Common Causes of Restarting Pods

In most production environments, restarting pods usually fall into one of the following categories.

CrashLoopBackOff

This is probably the most common Kubernetes error.

A CrashLoopBackOff doesn't tell us what failed—it simply tells us Kubernetes keeps restarting the container because it exits shortly after starting.

Common reasons include:

  • Application exceptions

  • Missing environment variables

  • Database connection failures

  • Invalid configuration

  • Missing ConfigMaps

  • Missing Secrets

  • Startup script failures

For example:

Error: Database connection refused

The application exits.

Kubernetes restarts it.

The application crashes again.

Eventually Kubernetes delays each restart attempt and reports CrashLoopBackOff.

It's important to remember that CrashLoopBackOff is a symptom, not the root cause.

The real reason is almost always available in the application logs.

OOMKilled

Another common reason for pod restarts is an out-of-memory condition.

If the application consumes more memory than allowed, the Linux kernel terminates the container.

We can verify this using:

kubectl describe pod <pod-name>

Look for something similar to:

Last State:
Terminated

Reason:
OOMKilled

Typical causes include:

  • Memory leaks

  • Processing large files

  • Image or PDF conversion

  • Loading large datasets into memory

  • Incorrect resource limits

Many teams immediately increase the memory limit and redeploy the application.

While that may stop the restarts temporarily, it's always worth understanding why the application is consuming excessive memory before increasing resources.

Readiness Probe Failures

A readiness probe tells Kubernetes whether the application is ready to receive traffic.

Example:

readinessProbe:
  httpGet:
    path: /health
    port: 8080

If the readiness probe fails:

  • The pod continues running.

  • Kubernetes stops routing traffic to it.

Common reasons include:

  • Wrong endpoint

  • Wrong port

  • Slow application startup

  • Database not yet available

  • Dependent services still starting

Readiness failures usually indicate that the application isn't fully initialised yet.

Liveness Probe Failures

Unlike readiness probes, liveness probes determine whether the application is still healthy.

If a liveness probe keeps failing, Kubernetes assumes the application is unhealthy and restarts it automatically.

Example:

livenessProbe:
  httpGet:
    path: /health
    port: 8080

One common mistake is configuring the liveness probe too aggressively.

For example, if an application requires 60 seconds to initialise but the liveness probe starts after only 15 seconds, Kubernetes may repeatedly kill the container before it has a chance to finish starting.

ImagePullBackOff

Sometimes the container never starts because Kubernetes cannot download the container image.

Possible reasons include:

  • Incorrect image name

  • Wrong image tag

  • Private registry authentication failure

  • Registry connectivity issues

  • Image does not exist

The quickest way to investigate is:

kubectl describe pod <pod-name>

Again, the Events section usually contains the exact reason for the image pull failure.

Configuration Problems

Configuration errors are another common cause of restarting pods.

Examples include:

  • Missing ConfigMaps

  • Missing Secrets

  • Incorrect environment variables

  • Invalid file paths

  • Missing certificates

  • Invalid application configuration

Even a simple typo in an environment variable can prevent an application from starting successfully.

Step 4: Check Resource Usage

Sometimes the problem isn't the application itself.

The Kubernetes node may be under heavy resource pressure.

We can check resource usage using:

kubectl top pod

kubectl top node

These commands help identify:

  • High CPU usage

  • High memory usage

  • Resource spikes

  • Node pressure

If these commands don't work, ensure the Kubernetes Metrics Server is installed.

Step 5: Review Cluster Events

Cluster events often provide additional context that application logs cannot.

Run:

kubectl get events --sort-by=.metadata.creationTimestamp

Look for messages related to:

  • Failed scheduling

  • Failed mounts

  • Failed image pulls

  • Probe failures

  • Resource exhaustion

Events frequently tell the complete story of what happened before the pod entered its restart cycle.

Step 6: Review Recent Deployments

If the application was working yesterday but started restarting after a deployment, don't ignore the possibility that the latest release introduced the issue.

Review the deployment history:

kubectl rollout history deployment payment-api

If required, roll back to the previous version:

kubectl rollout undo deployment payment-api

Rolling back isn't the final solution, but it can restore service quickly while we continue investigating the root cause.

A Simple Troubleshooting Checklist

Whenever we encounter a restarting pod, following a consistent process saves time and avoids unnecessary guesswork.

  1. Check the pod status.

  2. Describe the pod.

  3. Review the current logs.

  4. Review the previous logs.

  5. Check the Events section.

  6. Verify CPU and memory usage.

  7. Inspect readiness and liveness probes.

  8. Check for image pull errors.

  9. Review recent deployments.

Following the same sequence every time helps us diagnose problems much faster.

Common Mistakes

Over the years, we've seen the same mistakes repeated in many Kubernetes environments.

  • Deleting the pod before collecting logs.

  • Ignoring the Events section.

  • Increasing memory without understanding the root cause.

  • Restarting deployments repeatedly instead of investigating.

  • Forgetting to use kubectl logs --previous.

  • Assuming Kubernetes is the problem when the application itself is crashing.

Avoiding these mistakes can significantly reduce troubleshooting time.

Kubernetes doesn't restart containers without a reason. Every restart is a symptom of an underlying issue, whether it's an application crash, resource exhaustion, configuration error, failed health check or infrastructure problem.

The goal isn't to memorise dozens of Kubernetes commands. Instead, we should develop a structured troubleshooting approach that helps us identify the root cause quickly and consistently.

The next time you see a pod stuck in CrashLoopBackOff, resist the temptation to delete it immediately. Start by checking the pod status, describing the pod, reviewing the logs and inspecting the Events section. In most cases, Kubernetes has already provided enough information to lead us to the solution.

Sunday, October 2, 2022

NodeJs Read XML file and Parse Data

Hello,

Few days back I worked on task to read XML file in NodeJs and parse the data and convert data to the JSON format. Here in this blog I am going explain how I did achieved this. 

Following is our sample XML file.

<root>

    <parent>

        <firstchild>First Child Content 1</firstchild>

        <secondchild>Second Child Content 1</secondchild>

    </parent>

    <parent>

        <firstchild>First Child Content 2</firstchild>

        <secondchild>Second Child Content 2</secondchild>

    </parent>

</root>

Step 1 : Install necessary NPM packages

First install following packages in your NodeJs application.

npm install --save xmldom

mpm install --save hashmap

Here xmldom is the package we are going to use to parse the XML data and hasmap package we are going to use to convert and save data in JSON format.

Step 2 : Import necessary packages in script

const { readFile } = require('fs/promises');

const xmldom = require('xmldom');

const HashMap = require('hashmap');

We are going to use readFile from fs/promises. Also we will use xmldom to parse the string data. xmldom is a A JavaScript implementation of W3C DOM for Node.js, Rhino and the browser. Fully compatible with W3C DOM level2; and some compatible with level3. Supports DOMParser and XMLSerializer interface such as in browser.

Step 3 : Read the XML file

const readXMLFile = async()=> {

var parser = new xmldom.DOMParser();

        const result = await readFile('./data.xml',"utf8");

        const dom = parser.parseFromString(result, 'text/xml');

}

Here we are using async function to read the XML file as we want to wait till XML file is completely finished. Also we created a xmldom.DOMParser which we will use to parse the string data of the file. This parser will give you all the TAGs of the XML file just like we access standard HTML tags. With this we can use getElementsByTagName to get XML tags.  

Step 4 : Convert XML data to hasmap

const readXMLFile = async()=> {

var parser = new xmldom.DOMParser();

        const dataMap = new HashMap();

        const result = await readFile('./data.xml',"utf8");

        const dom = parser.parseFromString(result, 'text/xml');

        var parentList = dom.getElementsByTagName("parent");

        for(var i =0; i < parentList.length; i++) {

        const parent = parentList[i];

            const firstchild = parent.getElementsByTagName("firstchild")[0].textContent;

        const secondchild = parent.getElementsByTagName("secondchild")[0].textContent;

            const parentMap = new HashMap();

        parentMap.set("firstchild", firstchild);

        parentMap.set("secondchild", secondchild);

            dataMap.set(i, parentMap);

        }

}


In the above function we got list of parents by getElementsByTagName method and then we are iterating  through it and accessing the child tags and getting it's text content. Then we are simply storing it in hasmap with unique key. 

Hope this helps you.