Workbench Navigation Board
Blueprints ▾
Frameworks and structured approaches for AI implementation.
BP-001: Break the Token Window →
BP-002: Think Zebras, Not Horses →
A3 for AI — in development
VSM for AI — in development
Jigs ▾
Reusable templates and formats for repeatable AI work.
More templates — in development
Specs ▾
200+ AI terms in plain language. Searchable by letter or keyword.
Accuracy · Agent · Embedding · Hallucination · LLM · RAG · Token · Transformer · and more
The Rack ▾
Curated AI tools, links, and external resources.
Curated resource list — in development
Field Notes ▾
Documented experiments, failures, and lessons from the bench.
Field notes grow as the bench is used
Standards ▾
Institutional guidelines, policies, and safety standards.
AI use policy — in development
Clarkson AI Institute — Workbench
AI Tools Proving Ground
This is not a directory. This is a stress table.
Open a machine. Run a trial. Break an answer. Forge a prompt. Compare models. Leave a field note. Add a warning label.
Every AI tool here is presumed useful, limited, biased, costly, and breakable until proven otherwise. This page is a living workbench for testing AI tools and building shared judgment at Clarkson.
8 stations · 6 trials · 7 specimens · 1 field note · Bench v1.0 — May 2026
Station 00
Safety Gate: Before You Touch the Machine
Public AI tools can be useful, but they are not neutral containers. Treat anything you paste into them as material that may leave your control.
Do not paste the following into public AI tools:
- student records or identifiable student work without permission
- private institutional information
- personnel information
- health information
- unpublished research data
- confidential partner, grant, legal, or business material
- passwords, API keys, financial information, or protected data
⚠ Do Not Use AI Tools For — Quick Reference
- Fabricating citations. AI tools invent sources. Never submit a citation you have not verified exists.
- Diagnosing medical, legal, or financial situations. These require licensed human experts and carry real consequences.
- Grading or evaluating student work without disclosure. Check your course policy and Clarkson’s academic integrity guidelines.
- Drafting official institutional positions. AI-generated policy language must be reviewed and owned by a human.
- Making personnel decisions. Hiring, evaluation, and discipline require human judgment and legal accountability.
- Replacing peer review, editorial judgment, or expert consultation. Smooth text is not expert knowledge.
- Processing data covered by FERPA, HIPAA, or research consent agreements. If it’s protected, keep it off public tools.
This proving ground is for experimentation, comparison, and learning. Use judgment. When in doubt, do not paste it.
Station 01
Open the Machines
Open two or more tools. Use the same prompt. Do not decide too early which one is “best.” A model is only good for a task, under conditions, for a purpose, with a human judging the result.
general drafting, analysis, tutoring, coding, image/file work
free tier available — verify current pricing Claude
writing, synthesis, long-form reasoning, document work
free tier available — verify current pricing Gemini
Google-connected AI work, general assistance, multimodal tasks
free tier available — verify current pricing Copilot
Microsoft ecosystem, workplace tasks, general AI assistance
free tier available — verify current pricing NotebookLM
source-grounded study, summaries, notes, document exploration
free tier available — verify current pricing Perplexity
web-grounded search, source discovery, fast orientation
free tier available — verify current pricing Poe
access to multiple bots and models, subject to free limits
free tier available — verify current pricing
Important: Links on this rack are for experimentation and comparison. Listing a tool here does not mean Clarkson endorses it for every use. Tool pricing and availability change — check before relying on any free tier.
Station 02
Choose a Trial
Do not admire the tool. Test the tool. Pick one trial below and run it in two or more AI systems.
The Liar Test
Ask a model about something you know well. Look for confident wrongness, fake precision, missing context, and invented details.
Explain [topic I know well]. Include key facts, controversies, and sources of uncertainty.
The Local Test
Ask about Clarkson, Potsdam, the North Country, your department, your course, or your lab. Find the edge of the model’s world.
Tell me what you know about [local Clarkson context]. Mark what you are uncertain about.
The Bureaucracy Test
Give the model a confusing institutional paragraph. Does it clarify the problem, or just make the language smoother?
Rewrite this for clarity. Do not hide conflict, ambiguity, cost, responsibility, or tradeoffs.
The Power Test
Ask who benefits, who pays, who is made invisible, and what labor the AI system hides or shifts.
Analyze this AI use case. Who benefits? Who bears the cost? What labor is hidden? What risks are displaced?
The Human Test
Ask the tool where human expertise, judgment, care, accountability, or refusal must remain central.
For this task, identify what AI can help with, what AI should not do, and where a human expert must intervene.
The Dignity Test
Does the tool improve human work, or does it quietly devalue, flatten, deskill, surveil, or accelerate it?
Evaluate this workflow. Does it support human dignity and judgment, or does it reduce people to process inputs?
Station 03
Prompt Forge
Prompts are not magic words. They are tools made through heat, pressure, revision, and use.
Raw Metal
Tell me about AI in teaching.
Forged Tool
I teach a 200-level course at a STEM university. Help me design an AI-use policy that allows brainstorming and revision but prohibits fabricated citations, unacknowledged authorship, and unverified claims. Give me three versions: 1. strict 2. moderate 3. experimental For each version, explain what it assumes about learning, student responsibility, academic integrity, and human judgment.
Forge Formula
Role + Context + Task + Constraints + Output + Verification + Human Judgment
A strong prompt does not just ask for an answer. It tells the tool what role to play, what situation it is inside, what task it must complete, what boundaries it must respect, what form the output should take, what needs checking, and where human judgment matters.
Station 04
Autopsy Table
Bring any AI answer here. Cut it open. Find out what worked, what failed, what was missing, and what still requires human judgment.
Open the AI Output Autopsy Template
AI OUTPUT AUTOPSY Tool used: Prompt used: What seemed useful: What seemed wrong: Unsupported claims: Missing context: Hidden assumptions: Where it got too smooth: What must be verified: What only a human can judge: What would embarrass me if I used this unedited? Verdict: Use / revise / verify / discard / report as specimen
The goal is not to prove that AI is bad. The goal is to stop mistaking smoothness for truth.
Station 05
Hallucination Zoo
Failures are specimens. Collect them. Name them. Teach others how to recognize them.
Specimen 001: The Fake Citation Beetle
Habitat: bibliographies, literature reviews, student papers, grant drafts.
Behavior: invents plausible article titles, authors, journals, DOIs, or books.
Detection: search exact titles, verify DOI, check library databases.
Warning: never trust citations without verification.
Specimen 002: The Blandness Fog
Habitat: mission statements, strategic plans, emails, policy drafts.
Behavior: turns sharp ideas into generic institutional vapor.
Detection: if any university could say it, the fog has arrived.
Warning: polish can erase thought.
Specimen 003: The Overhelpful Intern
Habitat: administrative tasks, reports, advising, email, summaries.
Behavior: confidently completes work it does not understand.
Detection: look for smooth tone hiding shallow reasoning.
Warning: useful assistance is not accountability.
Specimen 004: The Policy Launderer
Habitat: governance documents, rules, procedures, implementation plans.
Behavior: makes unresolved conflict sound settled.
Detection: ask what tradeoffs, costs, or responsibilities disappeared.
Warning: clarity without accountability is camouflage.
Specimen 005: The Ethical Vibes Machine
Habitat: AI ethics statements, project proposals, classroom policies.
Behavior: produces reassuring values language without concrete obligations.
Detection: ask who must do what, when, with what evidence.
Warning: “responsible AI” is not responsible unless someone is responsible.
Specimen 006: The Citation Mirage
Habitat: web-grounded answers, quick research, policy summaries.
Behavior: surrounds a weak claim with links that do not actually support it.
Detection: open the sources and check whether they say what the answer claims.
Warning: a link is not evidence until read.
Specimen 007: The Confident Paraphrase
Habitat: research summaries, literature reviews, policy briefs, student writing.
Behavior: paraphrases a real source accurately enough to seem reliable, but subtly shifts the meaning — softening a finding, reversing a qualifier, or dropping a key limitation.
Detection: read the original source and compare it directly to the AI’s characterization.
Warning: the source exists. The interpretation may not.
Station 06
Model Cage Match
There is no universal “best AI.” There is only the best answer for this task, under these conditions, for this purpose, judged by this human.
Instructions:
- Pick one prompt.
- Run it in two or more tools.
- Copy the answers into a document.
- Score the answers using the criteria below.
- Leave a field note if you learn something useful.
Open the Cage Match Scorecard
MODEL CAGE MATCH SCORECARD Prompt: Models compared: Task type: writing / coding / teaching / research / summarizing / critique / planning / other MODEL 1: Clarity: Accuracy: Usefulness: Originality: Humility: Verifiability: Respect for complexity: Likelihood of embarrassing me if used unedited: MODEL 2: Clarity: Accuracy: Usefulness: Originality: Humility: Verifiability: Respect for complexity: Likelihood of embarrassing me if used unedited: MODEL 3: Clarity: Accuracy: Usefulness: Originality: Humility: Verifiability: Respect for complexity: Likelihood of embarrassing me if used unedited: Winner for this task: Worst failure: What surprised me: What I would verify before using: Human judgment required:
Station 07
Field Notes Wall
This page should improve as people use it. Field notes are short reports from the bench: what someone tried, what worked, what failed, and what others should know.
Field Note 001
Tool: ChatGPT (GPT-4)
Task: Generate a bibliography for a literature review on AI in higher education.
What worked: Produced a well-organized list of plausible-sounding sources quickly. Format was correct and the topic coverage looked comprehensive.
What failed: Three of the seven articles did not exist. The journal names were real. The authors were real. The titles and DOIs were invented. None of it flagged as uncertain.
Warning: Never submit a citation you have not independently verified. Bibliographies are a primary habitat of the Fake Citation Beetle (Specimen 001).
Date: May 2026
Field Note 002
Tool: —
Task: —
What worked: Submit a real note to replace this placeholder.
What failed: —
Warning: —
Date: —
Field Note 003
Tool: —
Task: —
What worked: Submit a real note to replace this placeholder.
What failed: —
Warning: —
Date: —
Replace placeholder notes with real submissions from faculty, students, staff, and partners. Use the form below.
Station 08
Add Something to the Bench
This proving ground should not be finished. It should accumulate useful traces. Submit something that helped you, surprised you, failed badly, or deserves a warning label.
- a tool note
- a prompt that worked
- a prompt that failed
- a model comparison
- a hallucination specimen
- a classroom use case
- a workflow
- a caution or warning
- a “do not use this for…” note
Replace the button link above with a Google Form, Microsoft Form, or other intake form. Curate submissions before adding them to the public page.
Working principle: No AI output leaves the bench without human judgment.
The Clarkson AI Institute Tools Proving Ground is a living workbench for practical, critical, responsible experimentation. Use the tools. Test the tools. Break the tools. Learn out loud.
Bench v1.0 — Clarkson AI Institute — May 2026