System logs follow a long-tailed distribution: a few templates appear constantly while many operationally important templates appear only a handful of times. An empirical study on the widely used Loghub-2.0 benchmark found that rare log groups (fewer than five instances) account for nearly 20% of all templates but less than 0.01% of log messages.

Because frequent templates dominate benchmark metrics, evaluation results look optimistic while poor performance on rare events stays hidden; all evaluated parsers degraded substantially on these rare groups.

TAILOR addresses this by enriching rare log groups with template-consistent log messages before template inference. The added structural evidence helps distinguish static tokens from dynamic variables, improving parsing accuracy on rare groups by 19% over the strongest baseline while keeping competitive performance on complete datasets, and it generalizes across different LLM backbones without modifying their core architectures.

Rare log groups cannot be fixed by collecting more data — they are rare in production by definition.

Synthesizing template-consistent messages adds the structural variation needed to tell static tokens from dynamic variables.

The augmentation sits outside the parser, so any LLM-based parser benefits without architectural changes.

Benchmark averages hid the failure: only a rare-group breakdown exposed it.

When a long tail is rare by nature you cannot collect your way out: synthesize examples that preserve the structure to be inferred, and existing models improve without retraining.

The paper reports a 19% parsing-accuracy gain on rare log groups over the strongest baseline, competitive performance on complete datasets, and consistent improvements across different LLM backbones without modifying their core architectures.

FOLLOW THE EVIDENCE

The sources

  1. TAILOR: Template-Preserving Augmentation for Long-Tailed Log Parsing arxiv.org