Auditing Legal prompts in English
Writing legal prompts in English is a different discipline from writing English prompts and translating the output. English is the default training distribution for every frontier model, so baseline quality is highest — but that also means you compete with the world's loudest prompts. Specificity beats volume. The auditor on this page scores your prompt against four 2026 frontier models — GPT-5, Claude 4.6, Gemini 3 Ultra and DeepSeek V3.2 — and ranks them by V-Index (quality per dollar) for contracts, charters, litigation memos and regulatory filings.
Legal work is unforgiving of model error. A hallucinated citation, a missed performance obligation, a wrong incoterm — each one costs hours or money downstream. The cheapest model is rarely the most expensive; the model that hallucinates least on your specific workload is. Our V-Index methodology measures both, in English, against the actual prompt you intend to ship.
Four pillars of a high-V-Index legal prompt
- Pillar 1
Jurisdiction-first framing
Always state the governing law, court, and applicable statute up front. A prompt that begins with 'Under DIFC law…' produces a fundamentally different draft than one that opens with 'Generally speaking…' — the former forces the model to retrieve jurisdiction-specific precedent, the latter invites hallucinated case names.
- Pillar 2
Cite-or-decline instructions
Add an explicit 'cite the exact section number or refuse to answer' clause. Hallucinated case citations are the #1 source of sanctions in 2026 — the prompt itself is your first line of defence.
- Pillar 3
Party and capacity disambiguation
Name the parties, their capacities, and the transaction stage. 'Draft a clause' is useless; 'draft a force-majeure clause for Party A as seller in a Saudi Aramco JV, governed by DIFC law, post-signing amendment' is auditable.
- Pillar 4
Output format pinning
Specify clause numbering, defined-term capitalisation, and whether you want a redline. Without this, every model picks a different convention and your paralegal spends 40 minutes on cleanup.
Five mistakes that tank legal prompt quality
- 01Asking for 'a contract' instead of a specific clause — produces a generic template instead of usable drafting.
- 02Omitting governing law — the model defaults to US/Delaware and you get the wrong indemnity standard.
- 03Forgetting to specify the audience (judge, opposing counsel, client) — tone and citation depth depend on it.
- 04Pasting the full agreement when only one section matters — wastes tokens, dilutes the model's focus.
- 05Trusting cited cases without verification — every model still fabricates citations at non-zero rates in 2026.
Three example legal prompts to audit
Each version below progressively adds the constraints discussed above. Run them through the auditor and watch the V-Index move.
Draft a force-majeure clause for a Saudi-Aramco joint-venture agreement governed by DIFC law.
Draft a force-majeure clause for a Saudi-Aramco joint-venture agreement governed by DIFC law, and end with a one-sentence summary for the partner.
Draft a force-majeure clause for a Saudi-Aramco joint-venture agreement governed by DIFC law. List every assumption explicitly. Refuse to answer any sub-question you cannot support with a cited source.
Frequently asked questions
- Which model is best for legal prompts in English?
- There is no universal answer — it depends on whether you optimise for cost, quality, or hallucination rate on your specific workload. The auditor on this page ranks GPT-5, Claude 4.6, Gemini 3 Ultra and DeepSeek V3.2 by V-Index for your exact prompt in English. As a rule of thumb in 2026: Claude 4.6 leads on legal reasoning tasks, DeepSeek V3.2 wins on cost-per-quality, GPT-5 is the safest all-rounder.
- Does prompt language affect output quality?
- Yes — significantly. Prompts in English route to different attention patterns than English prompts, even when the underlying request is identical. English is the default training distribution for every frontier model, so baseline quality is highest — but that also means you compete with the world's loudest prompts. Specificity beats volume. For high-stakes legal work, audit in both languages and compare.
- Is the free tier enough for legal work?
- The free tier (3 anonymous audits + 5/day signed-in) is enough to validate a prompt template you'll reuse. For daily legal work — refining client-specific prompts, generating PDF audit reports, switching between English and Professional English — Pro at $99/year removes the limits.
- How is V-Index calculated?
- V-Index = curated quality score (1–10) ÷ input price per 1M tokens (USD). A higher V-Index means more quality per dollar. The quality score is task-weighted: a model that is excellent at reasoning but weak at extraction will score differently for a legal extraction prompt than for a legal reasoning prompt.
- Are model citations reliable?
- No. Every frontier model in 2026 still fabricates citations at a non-zero rate, including the most expensive ones. The mitigation is in the prompt: require the model to refuse rather than guess, and verify every citation manually before shipping. Our audit reports flag citation-heavy prompts with an explicit hallucination-risk score.