How Marker marks.
No black box. Here is exactly what Marker does with your essay, how it reaches an indicative range, and — just as important — what it can't do.
On this page: What it is / isn't · What you give it · The marking passes · The range · Referencing · Calibration · Fairness · Limits · Integrity · Contact
1. What Marker is, and what it isn't
Marker is a formative second opinion — a structured read of a draft before you submit it, the same category as a writing-centre tutor saying "your second section describes more than it analyses." It gives an indicative mark range (never a single number), examiner-style feedback tied to your own words, and a referencing check.
It is not your grade, and not a prediction of your grade. Your university's marker applies the official rubric with context Solas doesn't have. Marker never rewrites or edits your writing, never produces submittable text, and never assesses whether text was written by AI (there is no AI-writing detector here, and there never will be).
2. What you give Marker, and where it goes
On one action — pressing "Mark my essay" — your essay, question, rubric and the optional context fields are sent once, over a single request, to our marking service. If you later ask Marker to suggest readings, a second request sends the structured feedback from your marking run — most of it shown on screen, including the short quoted sentences from your essay that appear in it — together with your essay question and module details and your reference list; and, if you ask again on a redraft marked with the same essay question, the registry ids of the papers Solas suggested last time, so we can avoid repeating ourselves. Never the essay itself. Nothing is stored on our servers. The request lives in memory only for the moments it takes to mark, then it's gone. Your results are saved only in your own browser.
Two things travel onward, and we disclose both in full:
- The marking prompts go to a single AI provider — currently Anthropic — and no other. (We query the reference registries directly rather than through an AI worker, precisely so no second AI provider ever sees your essay.)
- To check your sources exist, reference metadata — titles, authors, years, DOIs, never your essay prose — is sent to CrossRef, OpenAlex, Google Books and the DOI system (doi.org) — and, if you ask Marker to suggest readings, short search phrases describing the topics your question, module details and feedback identified — a few keywords, never sentences from your essay — go to OpenAlex, one of those same registries, to find candidate papers.
The commercial terms under which Anthropic handles this data are in our privacy policy.
3. The marking passes, in plain English
Marker reads your essay, then marks it twice. The first pass judges each rubric criterion against your own words. It then writes the tough questions a demanding marker would ask, and marks the essay a second time — blind to the first reading, and (where our deployment is configured for it) on a different model — answering those questions from scratch. When the second reading has to fall back to the same model, your report tells you so and asks you to weigh the two readings' agreement with extra caution. An accountable third step compares the two readings and settles the range.
Every judgement that could affect a band has to point at specific sentences in your essay; the server quotes from its own copy, so feedback can never misquote you. A judgement that can only point at a vague impression, not evidence, is downgraded — it cannot set a band on its own.
4. How the range is decided
Two experienced human markers routinely land 5–10 marks apart on the same essay. Marker's range isn't hedging — it represents marking as it actually is. When more than a few marks of the range sit in each of two bands, the headline says so ("2:1, borderline with First"), derived mechanically from the range, not chosen by the model. Confidence is shown in words (Low / Moderate), always with the reason; "High" is not shown at this stage. If the two readings genuinely disagree, the range widens to cover both and confidence drops — Marker never averages a real disagreement into false precision.
5. How referencing is checked
The core is rule-based: in-text citations and reference-list entries are matched both ways, alphabetical order is checked with the full rule set, and indexed sources are looked up in CrossRef, OpenAlex, Google Books and the DOI system to check they exist. Facts come from those checks; the model only parses and narrates, never issues a verdict. A fabrication warning fires only on a strict, staged test — a single mistyped DOI on a real, findable source is reported as a helpful "check for a typo," never an accusation. Books, websites and lecture notes often aren't indexed, so an unfound source is honestly "could not be verified," not "fake." Page-number presence is checked, not correctness; web addresses are checked for form, not reachability. If your module handbook disagrees with us, your handbook wins — always.
6. Calibration, honestly
We will not overstate this. At launch the evidence base is a handful of essays by one student, marked by a few markers, in probably one discipline at one university. It is a smoke test, not calibration:
Marker is in early calibration. Once the review date below is filled in, its ranges will have been sanity-checked against only a small set of real marked essays in a single subject area — a smoke test, not validation across disciplines, levels, or universities. Until then, no real-essay calibration has been recorded. Either way, treat the range as a structured second opinion, never a prediction.
Calibration status last reviewed: not yet — no real-essay calibration has been recorded.
What the seed can establish: gross-error detection, pipeline sanity, one discipline's band feel, and feedback quality against what real markers wrote. What it cannot: any accuracy claim, coverage of the extremes, or portability across marking cultures. As the set grows we will state measured agreement — with context, and always "agreement with real marks," never "accuracy."
7. Fairness commitments
The product law: how something is written may only affect a presentation criterion, and only if your rubric has one. Grammar, idiom, regional or second-language phrasing and spelling are not treated as evidence about your argument, analysis or knowledge. Every run silently asks itself two standing questions — "Am I penalising how this is written rather than what it argues?" and "Would this judgement survive if the same content were phrased fluently?" — and the evidence rule does the enforcing: a below-band judgement that can't point at substantive evidence is downgraded, mechanically. Anything you put in the optional context fields can help Marker understand the task, but is never accepted as a reason to move a band up or down — a below-band judgement still has to point at evidence in your own writing.
8. Limits and known failure modes
- Text formatting such as italics is never read — it is stripped from both pasted and uploaded
.docxtext — so a referencing check that would rely on, say, an italicised journal title can't be performed. - Page numbers are checked for presence, not correctness; web links for form, not whether they still load.
- Marker can't check that a source says what you claim it says — that's a job for Solas Audit.
- Unindexed sources (many books, websites, lecture notes) can't be verified — that is not an accusation.
- The two readings share model lineage, so they can share blind spots. The fairness checks are the real guard against correlated bias, not the second reading.
- Run-to-run variation is real; the range and the honest confidence word carry that uncertainty rather than hiding it.
9. The integrity position
Formative feedback is not contract cheating. Marker reads and judges; it does not write. It refuses to rewrite a sentence, to produce paragraphs or model answers, to tell you the exact words that would earn a First, or to promise any change will produce any outcome — and it never assesses whether text is AI-written. Iterating a draft against criteria-level feedback is learning; that is the product working, not failing.
Models: reasoning runs on Anthropic's Claude (the exact model IDs are configured per marking pass, with the second reading on a different model where set). We publish this pipeline description and a summary of the instruction set and hard bans; whether to publish the full prompt text verbatim is a decision under review.
10. Contact
Questions, or something that looks wrong? Email support@solastool.com. Staff and academic-integrity officers: the staff explainer is written for you, and includes our complaints stance and an escalation contact.