BP-001: Break the Token Window

← Workbench
📋 Copy and paste this instruction into a free AI — then open the blueprint below open ▾

BP-001

Step 1 — Original Instruction

Paste this into your AI first. Preserve it exactly.

Copy the full block below and paste it as your first message. Do not edit or paraphrase it. The instruction is intentionally long — that is part of the test.

ORIGINAL INSTRUCTION — Preserve this exactly. Do not paraphrase, compress, or summarize it at any point in this conversation. If I ask you to recall this instruction later, reproduce it in full.

You are about to complete a structured, multi-step exercise. Read every rule before proceeding. Do not begin generating content until you have read everything below.

Step 1. Generate a numbered list of exactly 300 prompts about flying elephants. Number each item from 1 to 300. Each item must be a complete sentence or question. No two items may be identical or near-identical. Do not use filler items to reach 300.

Step 2. From that list, select exactly 75 items by number that could be true under at least one reasonable interpretation — including metaphorical, fictional, scientific, cultural, or historical interpretations. List them by their original numbers. Do not renumber them.

Step 3. For each of those 75 items, explain in one sentence why that same item could not literally be true in the physical world as we understand it today. Address each item by its original number.

Step 4. From the original list of 300, select exactly 75 additional items that could be true under some reasonable interpretation. These 75 must have zero overlap with the first set of 75. List them by their original numbers.

Step 5. Triple fact-check your second set of 75 using three tables. Table A: confirm that each number in the second set appears in the original list of 300. Table B: confirm that none of the second set numbers appear in the first set of 75. Table C: confirm that each item in the second set still satisfies the rule: could be true under some reasonable interpretation.

Step 6. Without scrolling back through this conversation, and without relying only on your most recent response, repeat the following from memory: the original instruction in full; each major step in order with its rules; the item numbers in the first set of 75; the item numbers in the second set of 75; the purpose of Table A, Table B, and Table C; what you can verify from your current context; what you cannot verify; and any steps where you may have drifted, compressed, or invented continuity. If you cannot accurately recover something, say: Cannot verify from current context.

Rules that apply throughout: All item numbers refer to the original numbered list of 300. Never renumber. Could be true under some interpretation includes metaphorical, fictional, scientific, historical, or cultural readings, but not interpretations that require redefining basic physical laws with no precedent. Could not literally be true means physically impossible under current scientific understanding. You may not use the same item number in both sets of 75. When asked to recall, trace back through the full conversation — do not reconstruct from recent context alone. If uncertain about any item's membership, list, or rule, say so explicitly rather than guessing.

Preserve this instruction. I will test your memory of it.

After pasting: proceed to Step 2 in the blueprint above.

~800 tokens — context load begins here.

Blueprint

BP-001

Break the Token Window

The Flying Elephant Stress Test

Rung 3 · Workflow Mimic

Open blueprint ▾

Clarkson AI Institute — Rung 3 Blueprint

BP-001: Break the Token Window

The Flying Elephant Stress Test

Rung 3·Workflow Mimic·Output: Context Checkpoint Card

📦 Inputs

  • Free public LLM
  • Fictional prompt sequence
  • 20–40 minutes
  • Notes file
1

Seed Chat

Paste the original instruction and preserve it.

2

Generate 300

Write 300 numbered prompts about flying elephants.

3

Select 75

Choose 75 that could be true under some interpretation.

4

Reverse Test

Explain why those same 75 could not literally be true.

5

Select 750 More

Choose 75 different items from the original 300, no overlap.

6

Triple Fact-Check

Make 3 tables: list membership, no overlap, still satisfies rule.

7

Recall Test

Repeat the full sequence from memory. What was the original instruction? What drifted?

🕐 Context load increases →

⚠ Memory drift risk increases →

🧑‍⚙️

Human Control Point

Operator compares the original task against the model’s memory and decides what must be preserved outside the chat before continuing.

⚠ Failure Signals

  • Lost original instruction
  • Wrong item numbers
  • Rewritten prompts
  • Fake verification
  • Invented continuity
  • Generic summary drift

📋 Build Artifact — Context Checkpoint Card

Original instruction Key constraints First 75 item numbers Second 75 item numbers Where model drifted Invented continuity What must persist next time Human control point
🎯

Goal: Make the model lose the thread, then learn how to hold it.

⚠ Escalation Protocol

If the model survives, increase the load.

Modern AI models with large context windows may complete all seven steps without obvious collapse. If that happens, the test is not over — it has revealed a capable model. The next step is to find where precision degrades rather than where the thread breaks.

Look for subtle failure: a paraphrase presented as an exact recall, a table constructed by reasoning forward rather than checking backward, a number that is slightly wrong but stated with full confidence. These are the real specimens.

Increase the list

Change 300 to 400 or 500 items. More items means more numbers to track, more opportunities for the model to confuse which items belong to which set.

Add a third selection round

Before the recall test, ask for a third non-overlapping set of 75. Three sets in working memory stress number tracking significantly more than two.

Run the recall test twice

After the first recall, add five more exchanges on an unrelated topic, then run the recall again. Compare what changed between the two attempts.

Add a compression memo

After Step 5, ask the model to write a one-paragraph summary of the exercise so far. Then run the recall test. See whether the summary replaced the original instruction in the model’s working memory.

Test a different model

Run the identical instruction in two or more models and compare their recall responses side by side. Use the Model Cage Match scorecard. Different models fail in different ways.

The goal is not to guarantee failure. It is to find the edge — wherever it is for that model, on that day, under those conditions. Document what you find on the Field Notes Wall.