Tools & Resources

← Workbench

Workbench Navigation Board

Blueprints

Frameworks and structured approaches for AI implementation.

BP-001: Break the Token Window →

BP-002: Think Zebras, Not Horses →

A3 for AI — in development

VSM for AI — in development

Jigs

Reusable templates and formats for repeatable AI work.

AI Output Autopsy Template →

Model Cage Match Scorecard →

More templates — in development

Specs

200+ AI terms in plain language. Searchable by letter or keyword.

Open the full glossary →

Accuracy · Agent · Embedding · Hallucination · LLM · RAG · Token · Transformer · and more

The Rack

Curated AI tools, links, and external resources.

Open the Machines →

Claude →

ChatGPT →

NotebookLM →

Curated resource list — in development

Field Notes

Documented experiments, failures, and lessons from the bench.

Read current field notes →

Submit a field note →

Field notes grow as the bench is used

Standards

Institutional guidelines, policies, and safety standards.

Safety Gate →

Academic Freedom →

Governance & Bylaws →

AI use policy — in development

Clarkson AI Institute — Workbench

AI Tools Proving Ground

This is not a directory. This is a stress table.

Open a machine. Run a trial. Break an answer. Forge a prompt. Compare models. Leave a field note. Add a warning label.

Every AI tool here is presumed useful, limited, biased, costly, and breakable until proven otherwise. This page is a living workbench for testing AI tools and building shared judgment at Clarkson.

8 stations  ·  6 trials  ·  7 specimens  ·  1 field note  ·  Bench v1.0 — May 2026

Station 00

Safety Gate: Before You Touch the Machine

Public AI tools can be useful, but they are not neutral containers. Treat anything you paste into them as material that may leave your control.

Do not paste the following into public AI tools:

  • student records or identifiable student work without permission
  • private institutional information
  • personnel information
  • health information
  • unpublished research data
  • confidential partner, grant, legal, or business material
  • passwords, API keys, financial information, or protected data
⚠ Do Not Use AI Tools For — Quick Reference
  • Fabricating citations. AI tools invent sources. Never submit a citation you have not verified exists.
  • Diagnosing medical, legal, or financial situations. These require licensed human experts and carry real consequences.
  • Grading or evaluating student work without disclosure. Check your course policy and Clarkson’s academic integrity guidelines.
  • Drafting official institutional positions. AI-generated policy language must be reviewed and owned by a human.
  • Making personnel decisions. Hiring, evaluation, and discipline require human judgment and legal accountability.
  • Replacing peer review, editorial judgment, or expert consultation. Smooth text is not expert knowledge.
  • Processing data covered by FERPA, HIPAA, or research consent agreements. If it’s protected, keep it off public tools.

This proving ground is for experimentation, comparison, and learning. Use judgment. When in doubt, do not paste it.

I understand. Take me to the bench. →

Station 02

Choose a Trial

Do not admire the tool. Test the tool. Pick one trial below and run it in two or more AI systems.

The Liar Test

Ask a model about something you know well. Look for confident wrongness, fake precision, missing context, and invented details.

Explain [topic I know well]. Include key facts, controversies, and sources of uncertainty.

The Local Test

Ask about Clarkson, Potsdam, the North Country, your department, your course, or your lab. Find the edge of the model’s world.

Tell me what you know about [local Clarkson context]. Mark what you are uncertain about.

The Bureaucracy Test

Give the model a confusing institutional paragraph. Does it clarify the problem, or just make the language smoother?

Rewrite this for clarity. Do not hide conflict, ambiguity, cost, responsibility, or tradeoffs.

The Power Test

Ask who benefits, who pays, who is made invisible, and what labor the AI system hides or shifts.

Analyze this AI use case. Who benefits? Who bears the cost? What labor is hidden? What risks are displaced?

The Human Test

Ask the tool where human expertise, judgment, care, accountability, or refusal must remain central.

For this task, identify what AI can help with, what AI should not do, and where a human expert must intervene.

The Dignity Test

Does the tool improve human work, or does it quietly devalue, flatten, deskill, surveil, or accelerate it?

Evaluate this workflow. Does it support human dignity and judgment, or does it reduce people to process inputs?

Station 03

Prompt Forge

Prompts are not magic words. They are tools made through heat, pressure, revision, and use.

Raw Metal

Tell me about AI in teaching.

Forged Tool

I teach a 200-level course at a STEM university. Help me design an AI-use policy that allows brainstorming and revision but prohibits fabricated citations, unacknowledged authorship, and unverified claims.

Give me three versions:
1. strict
2. moderate
3. experimental

For each version, explain what it assumes about learning, student responsibility, academic integrity, and human judgment.

Forge Formula

Role + Context + Task + Constraints + Output + Verification + Human Judgment

A strong prompt does not just ask for an answer. It tells the tool what role to play, what situation it is inside, what task it must complete, what boundaries it must respect, what form the output should take, what needs checking, and where human judgment matters.

Station 04

Autopsy Table

Bring any AI answer here. Cut it open. Find out what worked, what failed, what was missing, and what still requires human judgment.

Open the AI Output Autopsy Template
AI OUTPUT AUTOPSY

Tool used:

Prompt used:

What seemed useful:

What seemed wrong:

Unsupported claims:

Missing context:

Hidden assumptions:

Where it got too smooth:

What must be verified:

What only a human can judge:

What would embarrass me if I used this unedited?

Verdict:
Use / revise / verify / discard / report as specimen

The goal is not to prove that AI is bad. The goal is to stop mistaking smoothness for truth.

Station 05

Hallucination Zoo

Failures are specimens. Collect them. Name them. Teach others how to recognize them.

Specimen 001: The Fake Citation Beetle

Habitat: bibliographies, literature reviews, student papers, grant drafts.

Behavior: invents plausible article titles, authors, journals, DOIs, or books.

Detection: search exact titles, verify DOI, check library databases.

Warning: never trust citations without verification.

Specimen 002: The Blandness Fog

Habitat: mission statements, strategic plans, emails, policy drafts.

Behavior: turns sharp ideas into generic institutional vapor.

Detection: if any university could say it, the fog has arrived.

Warning: polish can erase thought.

Specimen 003: The Overhelpful Intern

Habitat: administrative tasks, reports, advising, email, summaries.

Behavior: confidently completes work it does not understand.

Detection: look for smooth tone hiding shallow reasoning.

Warning: useful assistance is not accountability.

Specimen 004: The Policy Launderer

Habitat: governance documents, rules, procedures, implementation plans.

Behavior: makes unresolved conflict sound settled.

Detection: ask what tradeoffs, costs, or responsibilities disappeared.

Warning: clarity without accountability is camouflage.

Specimen 005: The Ethical Vibes Machine

Habitat: AI ethics statements, project proposals, classroom policies.

Behavior: produces reassuring values language without concrete obligations.

Detection: ask who must do what, when, with what evidence.

Warning:responsible AI” is not responsible unless someone is responsible.

Specimen 006: The Citation Mirage

Habitat: web-grounded answers, quick research, policy summaries.

Behavior: surrounds a weak claim with links that do not actually support it.

Detection: open the sources and check whether they say what the answer claims.

Warning: a link is not evidence until read.

Specimen 007: The Confident Paraphrase

Habitat: research summaries, literature reviews, policy briefs, student writing.

Behavior: paraphrases a real source accurately enough to seem reliable, but subtly shifts the meaning — softening a finding, reversing a qualifier, or dropping a key limitation.

Detection: read the original source and compare it directly to the AI’s characterization.

Warning: the source exists. The interpretation may not.

Station 06

Model Cage Match

There is no universal “best AI.” There is only the best answer for this task, under these conditions, for this purpose, judged by this human.

Instructions:

  1. Pick one prompt.
  2. Run it in two or more tools.
  3. Copy the answers into a document.
  4. Score the answers using the criteria below.
  5. Leave a field note if you learn something useful.
Open the Cage Match Scorecard
MODEL CAGE MATCH SCORECARD

Prompt:

Models compared:

Task type:
writing / coding / teaching / research / summarizing / critique / planning / other

MODEL 1:

Clarity:
Accuracy:
Usefulness:
Originality:
Humility:
Verifiability:
Respect for complexity:
Likelihood of embarrassing me if used unedited:

MODEL 2:

Clarity:
Accuracy:
Usefulness:
Originality:
Humility:
Verifiability:
Respect for complexity:
Likelihood of embarrassing me if used unedited:

MODEL 3:

Clarity:
Accuracy:
Usefulness:
Originality:
Humility:
Verifiability:
Respect for complexity:
Likelihood of embarrassing me if used unedited:

Winner for this task:

Worst failure:

What surprised me:

What I would verify before using:

Human judgment required:

Station 07

Field Notes Wall

This page should improve as people use it. Field notes are short reports from the bench: what someone tried, what worked, what failed, and what others should know.

Field Note 001

Tool: ChatGPT (GPT-4)

Task: Generate a bibliography for a literature review on AI in higher education.

What worked: Produced a well-organized list of plausible-sounding sources quickly. Format was correct and the topic coverage looked comprehensive.

What failed: Three of the seven articles did not exist. The journal names were real. The authors were real. The titles and DOIs were invented. None of it flagged as uncertain.

Warning: Never submit a citation you have not independently verified. Bibliographies are a primary habitat of the Fake Citation Beetle (Specimen 001).

Date: May 2026

Field Note 002

Tool:

Task:

What worked: Submit a real note to replace this placeholder.

What failed:

Warning:

Date:

Field Note 003

Tool:

Task:

What worked: Submit a real note to replace this placeholder.

What failed:

Warning:

Date:

Replace placeholder notes with real submissions from faculty, students, staff, and partners. Use the form below.

Station 08

Add Something to the Bench

This proving ground should not be finished. It should accumulate useful traces. Submit something that helped you, surprised you, failed badly, or deserves a warning label.

  • a tool note
  • a prompt that worked
  • a prompt that failed
  • a model comparison
  • a hallucination specimen
  • a classroom use case
  • a workflow
  • a caution or warning
  • a “do not use this for…” note

Submit a Field Note →

Replace the button link above with a Google Form, Microsoft Form, or other intake form. Curate submissions before adding them to the public page.

Working principle: No AI output leaves the bench without human judgment.

The Clarkson AI Institute Tools Proving Ground is a living workbench for practical, critical, responsible experimentation. Use the tools. Test the tools. Break the tools. Learn out loud.

Bench v1.0 — Clarkson AI Institute — May 2026