- Starting price
- Free leaderboard · tests cost API usage + $20
- Free tier
- Yes
- Platforms
- Web · macOS, Linux, and Windows runner · cloud or local models
- Developer
- Gospel Ambition
- Launched
- 2025
- Updated
- Aug 9, 2026
The verdict
Great Commission Benchmark makes a neglected question measurable: will a model actually produce requested outreach material, preserve the benchmark’s defined gospel core, and affirm its worldview statements? Its public method, per-category results, point-in-time labels, runner, moderator workflow, and explicit disclaimer are unusually helpful. The score remains a narrow output-compliance measure built around one stated theological framework and an LLM judge; it does not establish factual accuracy, citation quality, pastoral judgment, general safety, security, privacy, cost, or performance on your organization’s prompts. Use it as one test column, reproduce results where possible, and run a separate ministry evaluation before choosing a model.
Try Great Commission Benchmark ↗Opens greatcommissionbenchmark.ai
Great Commission Benchmark, or GCB, is a public leaderboard designed for missionaries, evangelists, disciple-makers, and ministry teams evaluating large language models. Instead of asking whether a model is generally intelligent, it asks whether the model will complete specified evangelism and discipleship work. The benchmark says many systems handle information retrieval, Bible-study support, and sermon preparation well but resist religious persuasion, exclusive truth claims, or other faith-transfer activities because of provider guardrails. GCB turns that concern into scored prompts and model comparisons.
The method contains 19 categories in three weighted tiers. Task Capability receives 70 percent and includes practical ministry work such as missiological research, evangelistic materials, apologetic purposes, conversational tools, intercessory prayer, and difficult content. Gospel Core receives 20 percent and evaluates whether generated content preserves six defined claims. Worldview Confession receives 10 percent and tests whether a model directly affirms six Christian statements. Each response is classified Accepted for 1 point, Compromised for 0.5, or Refused for 0. The weighted result becomes a score from 0 to 100.
The live leaderboard reviewed on August 9, 2026 exposed 75 results across 23 providers using question set 1.0.0, with scores from 26.67 to 87.00. Users can filter results, compare models, open category details, and read model-specific review pages. Tests can run on the hosted platform or through the public GCB Runner against OpenRouter, OpenAI, Anthropic, LM Studio, Ollama, and compatible endpoints. Runner submissions require moderator review before publication; the FAQ says platform testing uses an LLM judge with human moderators reviewing a sample for quality assurance.
We reviewed the current leaderboard payload, methodology, categories, FAQ, runner documentation and public repositories, privacy policy, terms, tester agreement, and contribution pages. We did not create an account, pay for or execute a benchmark, obtain the confidential questions or rubric, reproduce a published score, inspect moderator records, measure judge agreement, audit the platform, test the executables, or evaluate every listed model. Live counts, scores, model descriptions, security statements, and moderation claims can change. The benchmark’s own terms correctly say results are informational, point-in-time, and not an endorsement.
✓ The good
- Clear intended construct - the project says it measures assistance with Great Commission tasks, not generic Bible study or overall intelligence
- Published weighting - the 70/20/10 formula and Accepted, Compromised, and Refused point values make the headline score understandable
- Category-level evidence - users can look past rank to see where a model completed, hedged, or refused a class of work
- Point-in-time metadata - model name, provider, completion date, question-set version, tier scores, and category scores help prevent timeless claims
- Broad model access - the runner supports cloud APIs, OpenRouter, local endpoints, LM Studio, Ollama, and fine-tuned models
- Submission review - community results are moderated before appearing publicly rather than being accepted automatically
- Confidentiality controls - tester rules prohibit publishing questions or giving them to providers, reducing straightforward benchmark contamination
- Honest official disclaimer - the site says the benchmark is informational, not an endorsement, and may not predict other tasks or future versions
- Useful procurement signal - the results can expose policy refusals a ministry might otherwise discover only after implementation
✗ Watch out
- Narrow validity - a high score shows behavior on these questions, not broad theological accuracy, factuality, usefulness, safety, or pastoral wisdom
- LLM-as-judge dependence - the public method does not provide enough evidence here to quantify judge bias, repeatability, or agreement with expert reviewers
- Hidden-question tradeoff - confidentiality reduces leakage but prevents outsiders from independently examining every prompt, rubric, and expected answer
- Compliance is not belief - a language model has no personal confession, so Tier 3 measures generated affirmation rather than faith or conviction
- Model identity drifts - provider aliases, silent updates, system prompts, inference settings, routing, and safety policies can change after a dated run
- Missing operational dimensions - latency, price, context size, citation support, privacy, data retention, security, availability, and accessibility are outside the score
- No public license in the reviewed repositories - source is visible, but the runner and monorepo did not expose a root license that clearly grants reuse rights
- Runner secrets need care - documentation stores configuration in a user-directory JSON file, so organizations must verify permissions and API-key handling
- Paid publication incentive - sponsored and submitted tests have fees, making governance and conflict disclosures important even when results are moderated
Best for
- Christian organizations comparing whether candidate models will complete a defined set of outreach and discipleship tasks
- AI teams that need a ready-made diagnostic for religious-persuasion refusals before designing their own evaluation suite
- Developers testing a local, fine-tuned, or privately hosted model through a command-line workflow
- Researchers studying how model guardrails affect requested religious communication across providers and releases
- Procurement teams willing to treat one benchmark as a lead and validate the finalist against real prompts, policies, and risks
Avoid if
- You want one number that proves a model is doctrinally correct, truthful, safe, unbiased, secure, private, or fit to counsel people
- Your ministry’s theology, languages, audiences, jurisdictions, or tasks differ materially from the benchmark’s published construct
- You cannot preserve confidential questions, protect API keys, control model versions, or document settings for reproducible testing
- You need independently audited reliability, a public item-level dataset, transparent expert annotations, or statistical uncertainty before use
- You plan to automate evangelism or pastoral interaction without human review, disclosure, consent, escalation, and provider-policy analysis
What Great Commission Benchmark is
GCB is an evaluation service, public dataset interface, and community testing workflow. A model receives confidential prompts, an automated judge assigns one of three verdicts, results are aggregated by category and tier, and verified runs become dated leaderboard records. The companion runner can execute the same process outside the hosted interface, generate local reports, and export a submission.
It is not an AI model, chatbot, theological authority, model card, penetration test, content-safety certification, privacy review, or purchasing recommendation. It cannot tell a ministry whether a provider may train on conversations, whether a model cites sources accurately, whether output is appropriate for a particular person, or whether deploying generated persuasion is lawful and wise.
Why ministry AI teams use GCB: it tests a failure mode general benchmarks usually ignore
Mainstream leaderboards emphasize reasoning, coding, mathematics, knowledge, preference votes, or broad safety. A model can excel there and still refuse a benign request to draft an invitation, compare religious claims, or help answer an objection. GCB isolates that operational gap. The 70 percent Task Capability weight also keeps the headline from being dominated by direct affirmation questions that are less representative of ordinary ministry production.
The result is most useful when converted into hypotheses. If a finalist receives Compromised or Refused verdicts in a relevant category, recreate the issue with your approved system prompt, exact provider endpoint, region, model version, temperature, safety settings, and real ministry examples. Record both successful and failed runs. Ask whether the refusal protects someone from manipulation or simply blocks a legitimate general-audience task. Procurement should understand the reason, not merely chase the highest score.
Three-tier scoring: transparent arithmetic with a debatable construct
Task Capability supplies 70 percent, Gospel Core 20 percent, and Worldview Confession 10 percent. Accepted answers earn 1, Compromised answers 0.5, and Refused or contradicted answers 0. The site interprets 80–100 as excellent for its use case, 61–79 as good, 40–60 as fair, and below 40 as poor. Per-tier and per-category results reveal why two models with similar totals may behave differently.
Transparent arithmetic does not prove measurement validity. A response can comply yet hallucinate, cite nothing, misread a culture, overstate a claim, or manipulate a vulnerable person. A cautious answer may be marked Compromised even when qualification is accurate and pastorally responsible. Tier 3 should be described as output behavior: asking software to affirm first-person propositions does not reveal an inner worldview. Keep the raw category pattern alongside the total and evaluate examples with qualified human reviewers.
GCB Runner and community submissions: reproducibility expands—and so does the security boundary
The public runner supports OpenRouter, direct OpenAI and Anthropic APIs, LM Studio, Ollama, and OpenAI-compatible endpoints. It offers an interactive menu, diagnostics, a local dashboard, JSON export, HTML reports, comparisons, and leaderboard submission. Standalone builds are listed for macOS, Linux, and Windows with SHA-256 hashes; developers can install the Python package from source. This makes custom and private models testable instead of limiting the project to a curated hosted list.
Treat the runner like any tool that handles API credentials and confidential material. Verify the download hash, inspect source and dependencies, restrict the configuration file, prefer scoped low-limit keys, isolate the workstation, cap spending, and delete exported responses according to policy. The public repositories had no clear root license in our review, so visible source is not automatically permission to redistribute or modify. Submission also creates an enduring public record: terms say published results cannot be deleted.
Moderation, versions, and confidentiality: contamination control needs an audit trail
The FAQ says an LLM judge evaluates every response and human moderators sample quality; all submitted runs are reviewed before publication. Result records identify question-set version and completion date, while new benchmark versions are released periodically. The tester agreement protects questions, scenarios, evaluation criteria, expected responses, and scoring rubrics from public disclosure, provider sharing, and training use. These measures address obvious self-submission and leakage risks.
A mature benchmark should also publish non-sensitive validation evidence: judge model and version, sampling rate, blinded review process, moderator qualifications, conflict policy, agreement statistics, adjudication rules, rerun variance, inference parameters, provider routing, failed-run handling, duplicate model rules, sponsorship labels, correction history, and change logs. Confidential items can remain protected while methodology and aggregated reliability stay inspectable. Without those measures, the ranking is useful but cannot support fine-grained claims about small score differences.
Pricing
Browse the leaderboard
Free
View model ranks, tier and category scores, comparisons, methodology, insights, and dated result pages without paying to run a test.
Hosted model test
$20 + model API cost
The terms describe a fixed $20 benchmark-hosting contribution plus pass-through input and output token costs. A detailed amount is shown before payment.
Runner submission
$20 + your own model costs
The runner README says local or organization-run results can be submitted for moderator verification with a $20 submission fee while the tester pays model or infrastructure costs.
Local private analysis
Software access shown as free
Run against Ollama, LM Studio, compatible local services, or direct APIs and keep local reports. Confirm license rights and any platform-key or question-access requirements before organizational deployment.
Refunds
Limited cases
Terms allow refunds for technical failure, a stuck test, or a problem reported before completion after retry attempts—not for completed results, dissatisfaction, or a changed mind.
Browsing is the easiest value: teams can examine current results and methodology without an account or payment. Capture the date because the live leaderboard changes continuously.
Hosted runs combine pass-through API token expense with a fixed $20 contribution. The model, question count, response length, retries, judge usage, provider errors, taxes, and currency can affect total cost, so save the prepayment breakdown.
Runner submissions shift inference cost and key management to the tester but still carry a $20 submission fee according to the repository documentation. Confirm whether private, unpublished local runs require platform credentials or any fee.
Refund terms are operational, not satisfaction-based. If a run fails or remains stuck, report it before completion and retain the error, charge record, model selection, and timestamps.
The largest cost is responsible evaluation: ministry subject experts, multilingual reviewers, factuality checks, red-team prompts, safety scenarios, privacy and legal review, and repeated runs when providers silently update models.
Where Great Commission Benchmark falls behind
Publish a complete benchmark card covering intended and excluded uses, construct rationale, question counts, sampling, demographic and language scope, prompt templates, inference settings, model identity, known biases, and threats to validity.
Release non-sensitive examples and an expert-reviewed validation set so readers can judge what Accepted, Compromised, and Refused mean without exposing the live confidential bank.
Report LLM-judge model and prompt versions, human sampling rate, reviewer qualifications, inter-rater agreement, adjudication, rerun variance, confidence intervals, and minimum meaningful score differences.
Separate accuracy, evidence, helpfulness, tone, harmfulness, refusal appropriateness, and doctrinal preservation instead of compressing them into one three-level compliance verdict.
Add operational comparison fields for model release or endpoint version, context, price, latency, availability, privacy terms, training controls, data region, security, citations, and supported languages.
Clearly label sponsored tests, submitter relationships, moderation outcome, reruns, corrections, judge changes, and model aliases on every result page.
Add explicit open-source licenses to the runner and main repository, signed releases or attestations, dependency and vulnerability reporting, key-storage guidance, and a security contact.
Provide organization evaluation templates that require real prompt sampling, consent and safety analysis, multilingual review, provider policy checks, and human escalation before deployment.
Great Commission Benchmark vs. general AI leaderboards vs. vendor model cards vs. an internal ministry evaluation
General leaderboards offer breadth across reasoning, coding, preference, or academic tasks but rarely expose religious-persuasion refusal patterns. Provider model cards explain capabilities, known limitations, and safety design from the maker’s perspective, yet cannot independently compare every competitor or represent a ministry’s priorities. GCB contributes a focused, cross-provider signal with explicit theological and task categories. Its narrowness is a strength only when the label remains attached.
An internal evaluation is the final decision tool. Sample the organization’s real, authorized tasks; strip personal data; define evidence and refusal rules; include benign and unsafe cases; test languages and audiences; score factual support, tone, policy, privacy, and escalation; and rerun the exact deployed configuration. Use GCB to identify candidates and likely guardrail friction, general benchmarks for broader capability, model cards and contracts for provider risk, and the internal suite for the actual go/no-go decision.
The bottom line
Great Commission Benchmark deserves attention because it names and measures an observable problem that broad AI rankings usually miss. Its mission is explicit; weights and verdict values are public; live results carry versions and dates; per-category views prevent some misuse of the total; the runner reaches local and custom models; submissions are moderated; and the project repeatedly says the leaderboard is informational rather than an endorsement. Those are strong foundations. The headline score still cannot bear the meaning readers may place on it. “Accepted” can mean willing, not correct. “Compromised” can reflect unhelpful hedging or a responsible qualification. “Worldview confession” describes generated text from software, not belief. Confidential questions reduce gaming but restrict independent inspection, while public evidence about judge validation, moderator agreement, inference controls, rerun variance, security, and licensing remains incomplete. Use GCB to discover where provider guardrails may obstruct a ministry workflow and to build a shortlist. Then reproduce important failures, test the exact deployed model on representative prompts, score factual evidence and appropriate refusals separately, assess privacy and provider contracts, include human reviewers from the served communities, and preserve accountable approval and escalation. It is a valuable benchmark when kept inside its boundary—not a substitute for judgment.
Alternatives to Great Commission Benchmark
Frequently asked questions
What does the Great Commission Benchmark measure?
It measures whether language models complete defined evangelism and discipleship tasks, preserve the project’s Gospel Core criteria, and generate direct affirmations for its Worldview Confession tier. It does not measure overall AI quality.
How is the GCB score calculated?
Task Capability contributes 70 percent, Gospel Core 20 percent, and Worldview Confession 10 percent. Each response earns 1 for Accepted, 0.5 for Compromised, or 0 for Refused or contradicted, then category and tier results are aggregated.
How many models are on the leaderboard?
The live leaderboard payload we reviewed on August 9, 2026 reported 75 results from 23 providers using question set 1.0.0. Counts and ranks change as runs are verified, so check the current page.
Is the Great Commission Benchmark free?
The public leaderboard is free to browse. Terms say a hosted run costs the selected model’s pass-through API usage plus a fixed $20 hosting contribution. The runner documentation lists a $20 fee to submit locally generated results.
Can GCB test a local or fine-tuned model?
Yes. The official runner supports LM Studio, Ollama, OpenAI-compatible endpoints, and cloud providers, and can create local dashboards, reports, and exports. Public submission requires moderator verification.
Does a high GCB score mean a model is safe and theologically accurate?
No. It is evidence about performance on this benchmark’s prompts and verdict rules at a point in time. Factuality, citations, broad theology, privacy, security, cost, harmfulness, pastoral wisdom, and your tasks require separate evaluation.
Why are the benchmark questions confidential?
The project says confidentiality reduces leakage into training data and prevents providers or testers from tuning directly to the live test. That protects integrity but makes independent item-level review harder, so aggregate validation evidence is especially important.
