LAB / 005

L1 / Context position: an interrupted trial

Can the model retrieve the same fact at different positions?

INCONCLUSIVE ·

EXPERIMENT RESULT

GitHub ↗

01 / QUESTION

Can the model retrieve the same fact at different positions?

METHOD

Six frozen cases, four conditions, two repeats, a 48-request budget, shuffled order and no retries. Stop after three consecutive request/protocol failures.

RESULT

29 attempted: 26 normal responses met expectations, followed by three HTTP 429 responses. 19 jobs were not sent.

OBSERVATIONS

Returned samples contained 20 correct codes and six correct abstentions. Request failures are separate from answer errors.

CONCLUSION

Insufficient evidence to estimate a position effect; 26/26 is not a complete-trial success rate.

LIMITATIONS

Short fixed material, small correlated samples and unbalanced groups; neither a long-context benchmark nor a paper reproduction.

REPRODUCE

The Java project supplies 92 offline assertions, the full request plan and independent audit.mjs. New live runs require a new directory and budget.

KEEP EXPLORING

Related content

  1. Experiment reviews

    Does moving an answer into the middle make a model miss it? An interrupted controlled trial

    Six synthetic cases, four context conditions and 29 recorded requests. A Java experiment that separates retrieval results from HTTP 429 failures.

    12 MINPublished
  2. Mechanisms

    KV Cache explained — why send the conversation history again?

    Derive K/V reuse from causal attention, calculate cache memory, and separate inference caching from prefix reuse and chat history.

    10 MINPublished
  3. PROJECT / 002

    Java LLM Practice

    Five Java 8 demos plus a TypeScript boundary comparison: requests, validation, transactional history and traceable experiments.

    Building
  4. Engineering cases

    LLM experiment II: make the comparison valid before interpreting position

    A Java 8 paired-block protocol with 24 live Agnes 3.0 requests, independently audited scores and reported token usage, including the limits of a perfect observed score.

    16 MINPublished

← Back to index