Adaptive learning platform
A learning operating system, not a question bank
The largest system we have built for ourselves. Not a question bank with a chatbot bolted on: a platform that models why a student loses marks rather than how many, decides what they should do next, and is built so that the parts of it powered by a model cannot take the rest down with them.
- Status
- Live
- Category
- Adaptive learning platform
- For
- Students preparing for the Russian ЕГЭ and ОГЭ examinations, and their tutors
- Languages
- Russian
Conventional preparation counts right and wrong answers. A student losing marks to arithmetic slips and a student who never understood the theory receive identical instruction, and only one of them is helped by it. The interesting problem is not delivering questions — it is inferring the cause of a failure and acting on it.
- 01
Every attempt is stored as an event that is never overwritten, so mastery of a topic is derived from the whole history rather than from the last answer
- 02
A lost mark is attributed to one of four causes — attention, arithmetic, theory or strategy — and the next work is chosen against the cause rather than the topic
- 03
An engine that scores each candidate task as (impact × probability) − (fatigue × cost) − risk, and is allowed to decide the student should do less
- 04
Below 30% accuracy over the last ten attempts it stops setting questions altogether and switches to a break or to theory — a rule the model does not get a vote on
- 05
A tutor that remembers the individual student between sessions, practises oral subjects by voice call, and reads a photographed handwritten solution to mark the reasoning rather than the final answer
- 06
Four groups of features — the tutor, the voice tutor, analytics and photo review, and grading — each run on their own supplier account, so a stalled request costs one of the four rather than all of them
Fifteen questions before the dashboard opens
A new account chooses its subjects and answers fifteen questions, and the dashboard is not shown until it has. The platform will not recommend work to a student it has not yet measured, so the first minute of a new user's time goes on a test rather than on anything built to keep them there. That is the wrong order for signing people up, and it is the order we kept.
What those fifteen questions produce is not a mark but a share of the loss assigned to each cause: 35% attention, 30% arithmetic, 20% theory and 15% strategy in the worked example our own documentation carries. Written answers are marked against the examination's published criteria rather than against a model's opinion of them, so the reading stays anchored to the standard the student will actually be marked by.
- Held at the subtopic
- Subject, then exam task topic, then subtopic — the estimate is kept at the narrowest of the three, because 'mathematics' is not an instruction anyone can act on.
- Moved by the gap, not the answer
- Each attempt shifts the estimate by a share of the distance between what was expected of that student on that subtopic and what happened.
- Slower where it already knows
- An answer on unfamiliar ground moves the estimate up to three times as far as the same answer on a subtopic already settled, so one lucky or careless attempt cannot rewrite a long-standing reading.
- No model has a hand in it
- The estimate is arithmetic over the recorded attempts; the AI layer reads the result rather than producing it.
The limits sit above the score, not inside it
The day's work is assembled in three tiers rather than served as one queue: what has to be done, taken from the weakest areas and the nearest deadlines; what would be good to do, chosen to balance the load; and what is optional. That structure is what lets the system prescribe less without leaving a blank screen — a shortened day still arrives as a day, with the work that remains marked as the part that matters.
Above the tiers sit three rules that are allowed to empty them. An estimate of how tired the session has made the student, once it passes four-fifths of its scale, returns rest instead of work. Three errors of the same kind in a row move the student onto another topic, on the reading that a fourth question of that kind is the system's failure rather than the student's. The accuracy floor listed among the capabilities above is the third. All three are read before the scoring output is, so the day's highest-scoring task can be discarded unread.
Written the other way — tiredness as a penalty inside the score, a poor run as a discount — a large enough expected gain would eventually outweigh them, and it would do so on precisely the day the student most needed to stop. A limit that can be outbid is a preference. The state that produced each decision is kept beside the decision, so the plan issued on a particular day can be reconstructed afterwards rather than argued about. For an owner running an AI system of their own, the question worth asking is not whether there is a limit but where it lives: inside the number being maximised, or above it.
What each feature does when the model does not answer
A supplier's allowance is counted against the account the keys were issued from, not against the keys themselves, so one request that hangs can spend the limit for everything else running under it. That is why the work sits on four separate accounts, two or three keys on each, ten in total. A key that trips its limit is set aside to cool and brought back rather than thrown away. Every reply is checked before a student sees it, so one that fails the check becomes another attempt rather than a confident wrong explanation. One supplier does most of the work, with two others held behind it.
None of that changes from feature to feature. What changes is what the student is doing at the moment the answer fails to arrive.
- Voice tutoring
- Fails immediately and says the line is busy, because the student is mid-sentence and an unexplained silence is worse than a refusal.
- The written tutor
- Asks again ten seconds later, which is short enough that the student is still on the same question.
- Analytics, plans and photo review
- Held in a queue for up to a minute, because nobody sits and watches a study plan being written.
- Grading and test generation
- Waits thirty seconds before trying again, since a mark is expected back that day rather than that second.
Where the theory runs out, there is nothing to send
A rule that stops practice has to send the student somewhere, and what it sends them to has to exist at the same level of detail as the diagnosis: not the chapter on quadratics, but the part of it covering the step that went wrong. Material reaches the platform in three stages. Questions are imported in bulk from a public catalogue, which is the simplest of the three and the least of what makes the material useful. Each is then gone over again and marked with the topics it tests, the mistakes it is known to attract and the hints that meet them. Theory is written last, into a fixed structure, section by section, so that a student can be sent to one part of a topic rather than handed the topic.
Across eleven subjects, 7,017 questions have been imported, 6,541 of them have been through the second stage, and 982 sections of theory have been written. The third number is the one that constrains the product, and it is carried per subject rather than averaged, because an average would absorb the only thing worth knowing about it. History has theory for all nineteen of its exam topics; mathematics has eight of eighteen, so on ten topics the stop rule has nothing to send to, and a student failing one of them is given the break rather than the explanation. The unglamorous half of a system like this is the material itself — having it complete, and described finely enough that one part of it can be sent against one kind of failure. The model was never the part in doubt.
A complaint is sorted before anyone reads it
Two things are recorded without anyone having to report them: three or more clicks inside a single second, which is a person stuck rather than a person browsing, and an error in the browser, caught as it happens rather than reconstructed afterwards from a description of it. Both are held ninety days and then discarded.
What does get reported is sorted before a person reads it, into one of eleven topics in the team's own channel, in a fixed order. An error in the code goes to the topic for errors in the code; failing that, a word in the report picks the topic; failing that, the page the person was on picks it; and failing all three, it lands in general bug reports. The last rule matters more than the first three, because it is the one that ensures nothing arrives unowned. Complaints about the material itself never enter that queue. They are sorted separately, into four kinds, because the platform starts from the assumption that some of what it teaches is wrong and gives that assumption its own desk. The four are not degrees of severity; they are different faults, with different owners and different repairs — which is what the product does with a student's mistake, and what the team does with its own.
- Wording of the question
- The fault is in how the question is asked rather than in what is being asked, and the repair is an edit to the text alone.
- The answer key
- Nothing about the question is wrong except what it accepts, so every attempt already marked against it was marked against the wrong answer.
- The theory itself
- The error sits in an explanation rather than in a question, which means a student was taught the wrong thing before being asked anything.
- Content that did not load
- The one kind of the four in which the material is correct and the software simply failed to deliver it.
4 causes
Attention, arithmetic, theory or strategy — the cause of a lost mark, not the score
< 30%
Accuracy over ten attempts at which practice stops — a rule the AI cannot override
1 photo
Of handwritten working — marked on the reasoning, not just the final answer
Two things carried straight into client work. Diagnosis before instruction: knowing which of several causes produced a failure is worth more than knowing that it failed. And a limit that the model cannot argue with — the product refuses to sell more practice to a student who is failing, and a bad day costs one feature rather than the platform. Both are decisions taken before the traffic arrives, because neither can be retrofitted under load.
Invariant
A learning operating system, not a question bank