Why we built it
The gap is not effort. It is the absence of a shared standard.
IELTS and TOEFL Writing correction in Korea rests almost entirely on an individual tutor’s experience. The problem is structural, not a lack of care.
Rater variance
Internalising the official descriptors takes time. When the yardstick differs, the same essay gets different feedback and the learner cannot tell what to trust.
A hard ceiling on throughput
A careful reader manages two or three essays an hour. Peak season means queues, and a school can only grow by hiring more tutors.
Feedback that stops before the cause
Marking an error right or wrong never separates a surface slip from a misread task. That is why the same mistake returns.
A history that never accumulates
If attempts and corrections are not recorded, progress cannot be measured, and a school cannot spot a student who is about to drop out.
So scoring, diagnosis, root-cause analysis, prescription and re-verification were built as one pipeline. Not to replace a human reader, but to take over the part that is hard for a human to repeat.
R&D programme
Five engines connecting diagnosis to prescription.
The service running today is Claude-based automatic scoring. The five modules below are the research programme that will deepen that pipeline. They are not implemented performance.
Claude-based automatic scoring against the official criteria, sentence-level diagnosis and generated prescriptions
- M1Development target
Multimodal answer structure recognition
Extracts paragraphs, sentences and structural markers from handwritten scripts and turns them into measurable signals.
- Models
- SAM2 · RT-DETR · DINOv2 · TrOCR
- Input
- Handwritten scan or typed manuscript, plus exam-type metadata
- Output
- Paragraph structure tags, sentence boundaries, structure completeness score, clean text
- M2Development target
Fine-grained diagnosis and anomaly detection
Locates passages that depart from the distribution of high-scoring answers and calibrates the score.
- Models
- PatchCore · Anomaly Transformer · LightGBM · BGE-M3
- Input
- Clean text, sentence embeddings, high-scoring corpus distribution
- Output
- Per-sentence anomaly scores, suspect spans, calibrated criterion scores
- M3Development target
Two-stage root cause classification with retrieval
Separates surface causes from deeper ones and cites the assessment descriptor behind each prescription.
- Models
- KoELECTRA · BGE-M3 · bge-reranker
- Input
- Suspect spans, full answer, descriptor and model-answer knowledge base
- Output
- Surface and deep cause labels, cited evidence, prescription priorities
- M4Development target
Churn prediction and intervention policy
Sees a learner drifting away early and adjusts when and how to intervene.
- Models
- CatBoost · Temporal Fusion Transformer · Contextual Bandit
- Input
- Attempt frequency, score trajectory, report views, prescription completion
- Output
- Churn risk score, four-week progress forecast, chosen intervention, admin alerts
- M5Development target
Generative prescription content
Produces model answers at the target band, practice items and step-by-step comments.
- Models
- Claude + LoRA · retrieval · structural constraint checking
- Input
- Cause–evidence–prescription triples, learner state, target band or score
- Output
- Tailored model answers, practice sets, correction comments, vocabulary and grammar cards
System architecture
The path one essay travels
A submitted answer flows from structure recognition through diagnosis, cause classification and prescription. Behaviour logs branch off into churn prediction. Learners get a report; administrators get an at-risk list.
UI
Learner web and mobile (sitting tests, reports, progress) and the institution console (cohort statistics, at-risk alerts)
Service
Authentication, subscriptions, test sessions, scoring orchestration, learning history, notifications
AI engine
Structure recognition, anomaly-based diagnosis, cause classification with retrieval, prescription generation, churn prediction
Data
Original answers, scoring results and embeddings, behavioural time series, descriptor and model-answer knowledge base
Quantitative targets
What we are aiming to reach, and by how much
These are targets for the research programme. They are not figures the current service achieves; verified results will be published as they arrive.
| Module | Metric | Target |
|---|---|---|
| M1 | Structure detection mAP@0.5 | ≥ 0.90 |
| M1 | OCR alignment CER | ≤ 5% |
| M2 | Score prediction MAE | ≤ ±0.3 band · ±2 points |
| M2 | Correlation with official scoring | Pearson r ≥ 0.85 |
| M3 | Cause classification macro-F1 | ≥ 0.82 |
| M3 | Evidence retrieval Recall@10 | ≥ 0.90 |
| M4 | Churn prediction AUC | ≥ 0.80 |
| M5 | Generated answers meeting format rules | ≥ 95% |
| M5 | Hallucination rate | ≤ 5% |
Intellectual property
A closed loop from diagnosis back to re-verification
The loop — diagnose, classify the cause, prescribe, verify the effect, diagnose again — is what we intend to protect. The three filings below are being prepared and have not been filed yet.
- 01Filing in preparation
Automatic writing diagnosis by aligning recognised answer structure with assessment criteria
Turning paragraph and sentence structure into measurable signals and mapping them onto criterion items
- 02Filing in preparation
Two-stage surface and deep cause classification with evidence-cited prescription generation
Sorting detected errors into cause layers and building prescriptions on cited descriptor text
- 03Filing in preparation
Verifying prescription effect through rewrite comparison to select the next training step
Comparing a rewrite against the original on the same prompt to measure effect and choose what to train next
Development roadmap
12 months
The schedule for the research programme.
Months 1–2
Data foundation
Cleaning real answers and scoring labels, assembling the descriptor and model-answer knowledge base
Months 3–4
Structure recognition
Fine-tuning segmentation and structural element detection on handwritten scripts
Months 5–6
Diagnosis and anomaly detection
Detecting anomalous spans against the high-scoring distribution and calibrating scores
Months 7–8
Cause classification and retrieval
Surface and deep cause classifiers with evidence retrieval and reranking
Months 9–10
Prescription and personalisation
Constraint-respecting generation plus churn prediction and intervention policy
Month 11
Beta operation
Field trials with partner institutions and individual test takers
Month 12
Refinement and launch
Applying review findings, verifying the quantitative targets, public release
Business model
Start with individuals, widen to institutions.
Individual subscriptions build the data; that data becomes the metrics an academy administrator actually needs.
- B2C
Individual subscription
A monthly plan per exam. The first mock test is free on subscribing.
- B2B
Institution licence
An admin console with cohort attainment, at-risk alerts and prescription distribution.
- B2G
Public and educational bodies
Extending into study-abroad preparation and institutional writing assessment.