Built a tiny tool after my own validation numbers kept lying to me — metrictrust (pip install)

Hello everyone. While fine-tuning small models, I kept running into the same trap three times over:

  • the model hit its best score mid-training, and I ended up keeping a worse final checkpoint
  • train accuracy was much higher than eval — memorizing, not learning — so the number I was looking at overstated real performance
  • the run oscillated, so the final number was really just one noisy sample, not a measurement

Loss/metric loggers show you the curve, but none of them tell you when the number at the end of it shouldn’t be trusted. So I built a small (~300 lines, zero required dependencies) tool that watches a training run and prints a plain-language verdict: trust / caution / do-not-trust, plus what to do about it.

Install:

pip install metrictrust

Use it with HuggingFace Trainer (one line):

python

from metrictrust import TrustCheckCallback

trainer = Trainer(
    ...,
    callbacks=[TrustCheckCallback(metric_key="eval_accuracy")],
)
trainer.train()   # verdict prints when training ends

Or with any training loop:

python

from metrictrust import TrustMonitor

mon = TrustMonitor(higher_is_better=True)
for epoch in range(epochs):
    train_acc, val_acc = train_one_epoch(...)
    mon.record(step=epoch, eval_metric=val_acc, train_metric=train_acc)

print(mon.verdict())

Runs fully local — nothing leaves your machine.

Example output from a real run of mine (final log said 44%; the model had actually hit 100% mid-training and degraded — I nearly drew the wrong conclusion from it):

[X] DO NOT TRUST
  final metric: 0.44   best: 1.00 (step 450)
  - Best was 1.00 at step 450; you ended at 0.44.
    You kept a worse checkpoint.
  - Run oscillated 3x — one run is an anecdote, not a measurement.

Genuinely asking: is this a problem you run into as well, or do you already have a good way to catch it?

GitHub: davidtokartoky · GitHub PyPI: metrictrust · PyPI