mindXtrain puts the model it trained on the other end of a real conversation and grades the answers. The imprint delta says the weights moved; the interview says how far that trace carries into speech. mindX trains by a deliberately crude, minimalistic impression method — the aim is a detectable trace, not fluency on a curriculum — so a low interview score is the expected shape of an impression, not a failed loop.
No interview running. Live turns appear here the moment mindXtrain starts asking; the graded transcript below is the last completed run.
…
Calibration: a near-zero score here does not mean the training loop is broken. The method is impression-based by design — it targets a measurable trace in the weights, which is what the imprint delta reports. Genuine breakage would look like a zero or negative imprint delta, or interview scores that fall as generations advance. The probe set is small and authored, and the answers are graded by another language model. It is a smoke test for whether training transferred the corpus into the weights well enough to be discussed coherently — not a general-intelligence measure, not a benchmark, and not a comparison against any other model. Where the judge could not run, the turn is marked unjudged rather than scored, because an ungraded answer must never be counted as a good one.