A 4.8 GB error_log: 15 million lines, 106 real problems

logs php wordpress woocommerce performance

A client’s WooCommerce store arrived as a single file: 4.79 GB of error_log, 15,007,367 lines, covering 05-Jun-2026 to 28-Sep-2026 — roughly 115 days of a PHP application complaining about itself. The ask was easy to say and awkward to execute: find every distinct error in it, then tell us which ones actually repeat.

The file does not open in an editor. It does not fit in a context window, or in fifty of them: by the usual four-characters-per-token rule of thumb, that log is well over a billion tokens. And reading it wholesale would be the wrong move even with infinite capacity, because more than 99% of those lines are the same handful of messages repeated. The problem was never reading speed. It was the shape of the work.

Why opening the file is the wrong first move

Everything that fails with large files fails for the same reason: the tool tries to hold the whole thing at once. An editor loads it into memory. A naive script calls .read() and allocates 4.8 GB. A chat session tries to stuff it into a prompt and burns its budget on the first percent.

Here is what that file looks like if you measure before you touch it:

  • 4.79 GB, 15,007,367 lines, single file, appended to for four months
  • 14,986,204 of those lines are complete log entries; the rest are continuation lines — SQL fragments and Stack trace: frames that belong to the entry above them
  • 0 orphan lines at the end of the parse, when the parser gets the entry grammar right
  • 1,225 distinct error texts once timestamps are stripped
  • 106 distinct error types after normalization

That last pair is the whole story. Fifteen million lines collapse to 1,225 unique messages, and then to 106 real problems. The median error type in that file accounts for about 140,000 lines.

One streaming pass, two dictionaries

The method that works is unglamorous and fits in a page. Stream the file line by line — memory then holds only the distinct messages, which is bounded by how many different errors exist (1,225), not by how many lines there are. Never hold the file; hold the findings.

Inside that single pass, keep two dictionaries:

  1. Verbatim. Key = the entry text with the timestamp removed, because the timestamp is the one part that always differs between two occurrences of the same error. Store the first full occurrence verbatim, so nothing is lost — the artifact file is a genuine inventory, not a paraphrase.
  2. Normalized. Key = the same text with the volatile parts masked: hex hashes, database and table names, site paths, line numbers in identifiers.

Then the one performance decision that mattered most: look up the dedup dictionary before running any normalization regexes. Logs are overwhelmingly repeats, so the common case should be a dictionary hit and two counter increments, not a regex chain. Checking first took a two-million-line dry run from 82 seconds to 5 seconds — a factor of sixteen. Doing the work in the other order is the classic mistake in this task class, and it looks harmless because the output is identical.

The full 15-million-line pass then took 41.8 seconds on a single-core machine with 1 GB of RAM, at roughly 20 MB of peak memory.

The part you cannot get wrong: normalization

Everything above is mechanical. The danger is the silent corruption of counts, and there are five traps I would warn anyone about:

  • Classify after stripping the timestamp, not before. A regex anchored at the start of the line expecting PHP Warning: matches literally nothing, because the line starts with [. Every entry then buckets into “other” and the report looks complete while being useless.
  • Mask quoted strings on word boundaries. A plain '[^']*' pattern treats an apostrophe in a contraction as an opening quote, pairs it with the next quote somewhere later in the line, and eats the actual message in between. In a log full of English prose this quietly deletes the evidence.
  • Never normalize the same string twice. The sentinel masks are not idempotent: a second pass over <SITE> turns it into ‘S’ and fragments your categories.
  • Mask volatile tokens deliberately. Long hex strings and temp-file names differ on every occurrence. Without masking, one logical error becomes hundreds of singleton “types” and the ranking is noise.
  • Cap the accumulator per entry. A single database error with an embedded query and a call chain can be enormous; without a cap, a few entries decide your memory profile.

And the rule that applies to every data reduction, not just logs: verify the totals against the source. The counts across both stages must sum to the number of parsed entries, the compressed inventory must pass its own integrity check, and the report must document every mask it applied so the numbers can be audited. A number you cannot re-derive is a number you cannot publish.

What the file was actually saying

With the noise collapsed, the story took one paragraph to tell. Of 15 million lines:

  • 11.2 million were PHP Deprecated notices — 74% of the entire file — coming from a small number of plugins calling functions that the PHP version in use has retired: a WooCommerce product-feed plugin accounted for about 669,000 of them, a table-rate-shipping plugin for roughly 665,000 plus another 490,000 across several call sites each.
  • 669,000 lines came from a translation plugin with debug logging left switched on. Someone had turned it on to diagnose something and never turned it off; a quarter of a million lines a month is what that costs.
  • 677,000 lines were a single line of configuration being loaded twice — Constant ABSPATH already defined — on every request for four months.
  • 1,724 fatal errors and 330 database errors hid underneath all of it. Those are the entries that correlate with actual downtime and lost orders, and in an unbounded file they are invisible: 0.01% of the lines, buried by fifteen million deprecation notices.

That distribution is the useful finding. Deprecations are not emergencies today; they are a schedule. A store running that many retired-function calls is a store that will break on the next PHP major version — the noise is a countdown, and the fix is a plugin inventory, not a bigger server. The 1,724 fatals are the opposite: rare, urgent, and previously unfindable.

The general recipe for a file too big to open

Nothing here is specific to logs. The same five moves handle a giant CSV, a dump, or a corpus:

  1. Measure before reading. File size, line count, then sample twelve random byte offsets and read thirty lines at each. You will learn every record format in the file without a single full scan — and you will not write a parser for a format that appears eleven times.
  2. Write one streaming pass. Read lines, never the file. Emit two artifacts: the complete inventory (compressed) and the grouped summary. Nothing else needs to exist.
  3. Dedupe before you transform. Repeats are the norm in real data; make the cheap check the first check.
  4. Make the reduction auditable. Document the masks, cap the accumulators, keep one verbatim sample per group so a human can confirm the grouping is honest.
  5. Verify the totals against the source, then keep the parser. In this project the parser is a committed script and the 4.8 GB log is in .gitignore — the repository holds 200 KB of findings and the method that produced them, not the input.

The deliverable of a triage like this is never a cleaned-up file. It is an index: every distinct error verbatim, with exact occurrence counts, first and last seen, and a ranked list short enough to act on. Per-occurrence enumeration would just reproduce the source log — which is what the file already was.

If you have a log, an export or a dump that has outgrown the tools you open it with, that is a solvable problem before it is a big one. Related reading from this blog: how we cut a 12 GB agent state database by 58% by changing the storage layout in the compact FTS migration, and how instrumentation answered a measurement question that had been settled by opinion in AnGo Scroll Analytics. Everything else is in the blog archive.

Figures come from a single measured triage run on one 4.79 GB log; client, site and plugin-level details are deliberately withheld. The streaming parser is a reusable script rather than a one-off.