Solas Marker · Transparency

How Marker marks.

No black box. Here is exactly what Marker does with your essay, how it reaches an indicative range, and — just as important — what it can't do.

On this page: What it is / isn't · What you give it · The marking passes · The range · Referencing · Calibration · Fairness · Limits · Integrity · Contact

1. What Marker is, and what it isn't

Marker is a formative second opinion — a structured read of a draft before you submit it, the same category as a writing-centre tutor saying "your second section describes more than it analyses." It gives an indicative mark range (never a single number), examiner-style feedback tied to your own words, and a referencing check.

It is not your grade, and not a prediction of your grade. Your university's marker applies the official rubric with context Solas doesn't have. Marker never rewrites or edits your writing, never produces submittable text, and never assesses whether text was written by AI (there is no AI-writing detector here, and there never will be).

2. What you give Marker, and where it goes

On one action — pressing "Mark my essay" — your essay, question, rubric and the optional context fields are sent once, over a single request, to our marking service. If you later ask Marker to suggest readings, a second request sends the structured feedback from your marking run — most of it shown on screen, including the short quoted sentences from your essay that appear in it — together with your essay question and module details and your reference list; and, if you ask again on a redraft marked with the same essay question, the registry ids of the papers Solas suggested last time, so we can avoid repeating ourselves. Never the essay itself. Nothing is stored on our servers. The request lives in memory only for the moments it takes to mark, then it's gone. Your results are saved only in your own browser.

Two things travel onward, and we disclose both in full:

The commercial terms under which Anthropic handles this data are in our privacy policy.

3. The marking passes, in plain English

Marker reads your essay, then marks it twice. The first pass judges each rubric criterion against your own words. It then writes the tough questions a demanding marker would ask, and marks the essay a second time — blind to the first reading, and (where our deployment is configured for it) on a different model — answering those questions from scratch. When the second reading has to fall back to the same model, your report tells you so and asks you to weigh the two readings' agreement with extra caution. An accountable third step compares the two readings and settles the range.

This is a second reading designed to resist anchoring — not an independent second marker. Both readings share model lineage and may share blind spots; that is exactly why the fairness checks below exist, and why we watch for the two readings agreeing too often.

Every judgement that could affect a band has to point at specific sentences in your essay; the server quotes from its own copy, so feedback can never misquote you. A judgement that can only point at a vague impression, not evidence, is downgraded — it cannot set a band on its own.

4. How the range is decided

Two experienced human markers routinely land 5–10 marks apart on the same essay. Marker's range isn't hedging — it represents marking as it actually is. When more than a few marks of the range sit in each of two bands, the headline says so ("2:1, borderline with First"), derived mechanically from the range, not chosen by the model. Confidence is shown in words (Low / Moderate), always with the reason; "High" is not shown at this stage. If the two readings genuinely disagree, the range widens to cover both and confidence drops — Marker never averages a real disagreement into false precision.

5. How referencing is checked

The core is rule-based: in-text citations and reference-list entries are matched both ways, alphabetical order is checked with the full rule set, and indexed sources are looked up in CrossRef, OpenAlex, Google Books and the DOI system to check they exist. Facts come from those checks; the model only parses and narrates, never issues a verdict. A fabrication warning fires only on a strict, staged test — a single mistyped DOI on a real, findable source is reported as a helpful "check for a typo," never an accusation. Books, websites and lecture notes often aren't indexed, so an unfound source is honestly "could not be verified," not "fake." Page-number presence is checked, not correctness; web addresses are checked for form, not reachability. If your module handbook disagrees with us, your handbook wins — always.

6. Calibration, honestly

We will not overstate this. At launch the evidence base is a handful of essays by one student, marked by a few markers, in probably one discipline at one university. It is a smoke test, not calibration:

Marker is in early calibration. Once the review date below is filled in, its ranges will have been sanity-checked against only a small set of real marked essays in a single subject area — a smoke test, not validation across disciplines, levels, or universities. Until then, no real-essay calibration has been recorded. Either way, treat the range as a structured second opinion, never a prediction.

Calibration status last reviewed: not yet — no real-essay calibration has been recorded.

What the seed can establish: gross-error detection, pipeline sanity, one discipline's band feel, and feedback quality against what real markers wrote. What it cannot: any accuracy claim, coverage of the extremes, or portability across marking cultures. As the set grows we will state measured agreement — with context, and always "agreement with real marks," never "accuracy."

7. Fairness commitments

The product law: how something is written may only affect a presentation criterion, and only if your rubric has one. Grammar, idiom, regional or second-language phrasing and spelling are not treated as evidence about your argument, analysis or knowledge. Every run silently asks itself two standing questions — "Am I penalising how this is written rather than what it argues?" and "Would this judgement survive if the same content were phrased fluently?" — and the evidence rule does the enforcing: a below-band judgement that can't point at substantive evidence is downgraded, mechanically. Anything you put in the optional context fields can help Marker understand the task, but is never accepted as a reason to move a band up or down — a below-band judgement still has to point at evidence in your own writing.

8. Limits and known failure modes

9. The integrity position

Formative feedback is not contract cheating. Marker reads and judges; it does not write. It refuses to rewrite a sentence, to produce paragraphs or model answers, to tell you the exact words that would earn a First, or to promise any change will produce any outcome — and it never assesses whether text is AI-written. Iterating a draft against criteria-level feedback is learning; that is the product working, not failing.

Models: reasoning runs on Anthropic's Claude (the exact model IDs are configured per marking pass, with the second reading on a different model where set). We publish this pipeline description and a summary of the instruction set and hard bans; whether to publish the full prompt text verbatim is a decision under review.

10. Contact

Questions, or something that looks wrong? Email support@solastool.com. Staff and academic-integrity officers: the staff explainer is written for you, and includes our complaints stance and an escalation contact.