Skip to main content

Why we built it

The gap is not effort. It is the absence of a shared standard.

IELTS and TOEFL Writing correction in Korea rests almost entirely on an individual tutor’s experience. The problem is structural, not a lack of care.

  • Rater variance

    Internalising the official descriptors takes time. When the yardstick differs, the same essay gets different feedback and the learner cannot tell what to trust.

  • A hard ceiling on throughput

    A careful reader manages two or three essays an hour. Peak season means queues, and a school can only grow by hiring more tutors.

  • Feedback that stops before the cause

    Marking an error right or wrong never separates a surface slip from a misread task. That is why the same mistake returns.

  • A history that never accumulates

    If attempts and corrections are not recorded, progress cannot be measured, and a school cannot spot a student who is about to drop out.

So scoring, diagnosis, root-cause analysis, prescription and re-verification were built as one pipeline. Not to replace a human reader, but to take over the part that is hard for a human to repeat.

R&D programme

Five engines connecting diagnosis to prescription.

The service running today is Claude-based automatic scoring. The five modules below are the research programme that will deepen that pipeline. They are not implemented performance.

Running today

Claude-based automatic scoring against the official criteria, sentence-level diagnosis and generated prescriptions

  • M1Development target

    Multimodal answer structure recognition

    Extracts paragraphs, sentences and structural markers from handwritten scripts and turns them into measurable signals.

    Models
    SAM2 · RT-DETR · DINOv2 · TrOCR
    Input
    Handwritten scan or typed manuscript, plus exam-type metadata
    Output
    Paragraph structure tags, sentence boundaries, structure completeness score, clean text
  • M2Development target

    Fine-grained diagnosis and anomaly detection

    Locates passages that depart from the distribution of high-scoring answers and calibrates the score.

    Models
    PatchCore · Anomaly Transformer · LightGBM · BGE-M3
    Input
    Clean text, sentence embeddings, high-scoring corpus distribution
    Output
    Per-sentence anomaly scores, suspect spans, calibrated criterion scores
  • M3Development target

    Two-stage root cause classification with retrieval

    Separates surface causes from deeper ones and cites the assessment descriptor behind each prescription.

    Models
    KoELECTRA · BGE-M3 · bge-reranker
    Input
    Suspect spans, full answer, descriptor and model-answer knowledge base
    Output
    Surface and deep cause labels, cited evidence, prescription priorities
  • M4Development target

    Churn prediction and intervention policy

    Sees a learner drifting away early and adjusts when and how to intervene.

    Models
    CatBoost · Temporal Fusion Transformer · Contextual Bandit
    Input
    Attempt frequency, score trajectory, report views, prescription completion
    Output
    Churn risk score, four-week progress forecast, chosen intervention, admin alerts
  • M5Development target

    Generative prescription content

    Produces model answers at the target band, practice items and step-by-step comments.

    Models
    Claude + LoRA · retrieval · structural constraint checking
    Input
    Cause–evidence–prescription triples, learner state, target band or score
    Output
    Tailored model answers, practice sets, correction comments, vocabulary and grammar cards

System architecture

The path one essay travels

A submitted answer flows from structure recognition through diagnosis, cause classification and prescription. Behaviour logs branch off into churn prediction. Learners get a report; administrators get an at-risk list.

Answer submitted
  1. UI

    Learner web and mobile (sitting tests, reports, progress) and the institution console (cohort statistics, at-risk alerts)

  2. Service

    Authentication, subscriptions, test sessions, scoring orchestration, learning history, notifications

  3. AI engine

    Structure recognition, anomaly-based diagnosis, cause classification with retrieval, prescription generation, churn prediction

  4. Data

    Original answers, scoring results and embeddings, behavioural time series, descriptor and model-answer knowledge base

Report · prescription · alerts

Quantitative targets

What we are aiming to reach, and by how much

These are targets for the research programme. They are not figures the current service achieves; verified results will be published as they arrive.

Development target
ModuleMetricTarget
M1Structure detection mAP@0.5≥ 0.90
M1OCR alignment CER≤ 5%
M2Score prediction MAE≤ ±0.3 band · ±2 points
M2Correlation with official scoringPearson r ≥ 0.85
M3Cause classification macro-F1≥ 0.82
M3Evidence retrieval Recall@10≥ 0.90
M4Churn prediction AUC≥ 0.80
M5Generated answers meeting format rules≥ 95%
M5Hallucination rate≤ 5%

Intellectual property

A closed loop from diagnosis back to re-verification

The loop — diagnose, classify the cause, prescribe, verify the effect, diagnose again — is what we intend to protect. The three filings below are being prepared and have not been filed yet.

  • 01Filing in preparation

    Automatic writing diagnosis by aligning recognised answer structure with assessment criteria

    Turning paragraph and sentence structure into measurable signals and mapping them onto criterion items

  • 02Filing in preparation

    Two-stage surface and deep cause classification with evidence-cited prescription generation

    Sorting detected errors into cause layers and building prescriptions on cited descriptor text

  • 03Filing in preparation

    Verifying prescription effect through rewrite comparison to select the next training step

    Comparing a rewrite against the original on the same prompt to measure effect and choose what to train next

Development roadmap

12 months

The schedule for the research programme.

  1. Months 1–2

    Data foundation

    Cleaning real answers and scoring labels, assembling the descriptor and model-answer knowledge base

  2. Months 3–4

    Structure recognition

    Fine-tuning segmentation and structural element detection on handwritten scripts

  3. Months 5–6

    Diagnosis and anomaly detection

    Detecting anomalous spans against the high-scoring distribution and calibrating scores

  4. Months 7–8

    Cause classification and retrieval

    Surface and deep cause classifiers with evidence retrieval and reranking

  5. Months 9–10

    Prescription and personalisation

    Constraint-respecting generation plus churn prediction and intervention policy

  6. Month 11

    Beta operation

    Field trials with partner institutions and individual test takers

  7. Month 12

    Refinement and launch

    Applying review findings, verifying the quantitative targets, public release

Business model

Start with individuals, widen to institutions.

Individual subscriptions build the data; that data becomes the metrics an academy administrator actually needs.

  • B2C

    Individual subscription

    A monthly plan per exam. The first mock test is free on subscribing.

  • B2B

    Institution licence

    An admin console with cohort attainment, at-risk alerts and prescription distribution.

  • B2G

    Public and educational bodies

    Extending into study-abroad preparation and institutional writing assessment.

Technology — Writing PT