The solution
The Outage Data Initiative Nationwide (ODIN) collects heterogeneous power-outage reports that must be transformed into standardized XML compliant with CIM IEC 61968-3. A small language model fine-tuned the usual way stalled at 16.20% overall accuracy.
The authors applied Minimum Risk Training (MRT), introduced for neural machine translation a decade ago: instead of relying on token-level maximum-likelihood objectives, training directly optimizes sequence-level evaluation metrics, so mistakes that break the target format are penalized during training.
With MRT, Qwen2.5-7B-Instruct reached 68.95% overall accuracy on the task — a more than fourfold jump — demonstrating the effectiveness of sequence-level optimization for domain-specific structured generation, in line with recent work showing renewed potential for risk-based optimization in modern language models.
Why it worked
Standardized XML either validates or does not, so a sequence-level metric captures the real objective better than token likelihood.
Direct metric optimization penalizes whole-output failures such as wrong fields or broken structure during training rather than after.
The technique predates modern LLMs but composes with them: no architecture change was needed.
A small 7B model suffices once the training objective matches the judged metric.
What can be applied
When output must satisfy a checkable format, make the check itself the training signal; a decade-old sequence-level trick can beat prompt-level fixes for narrow structured generation.
Aftermath
The paper reports the 16.20%-to-68.95% jump on ODIN's standardization task and frames it as evidence of renewed potential for risk-based optimization in modern language models.
FOLLOW THE EVIDENCE