Why does my Python script run out of memory processing a large file?
Almost always because the whole file is being loaded into memory at once. Reading a file with a method that returns its entire contents, or building a list of every line before processing, requires memory proportional to the file size. A two gigabyte file needs at least two gigabytes, usually considerably more once each line becomes a separate string object.
The fix is to process the file as a stream rather than as a value. Iterating over a file object yields one line at a time and keeps only that line in memory, so the footprint stays constant regardless of file size. This single change resolves most cases.
The same principle applies further along the pipeline, and this is where people often reintroduce the problem after fixing the read. Building a list of results defeats the purpose, because the list grows to match the input. Generators let each stage yield items as they are produced, so a chain of transformations still processes one record at a time. Any function returning a list can usually become a generator by yielding rather than appending.
Structured formats need format-aware streaming. Parsing an entire JSON document requires the whole thing in memory by design, which is why line-delimited JSON exists. Data analysis libraries generally accept a chunk size parameter that returns an iterator of pieces rather than one object.
Two related causes are worth checking. Accidentally holding references, for instance appending every processed record to a list for a final count, keeps everything alive. And reading a file in text mode with the wrong encoding can silently produce far more objects than expected.
If the algorithm genuinely needs all the data at once, the answer is usually a database or a chunked approach rather than more memory.