Brandon K / research notes

Prompting × training · 19 September 2026

Same behavior.
Different route?

An instruction can change what a language model does. Training can too. If both produce the same behavior, do they change the model’s internal computation in similar ways?

Work in progress · steering calibration complete; held-out evaluation pending

Measured · 64 calibration items · one training seed

The routes pass the matching rule.

Epoch 1 is the earliest checkpoint passing all three gates. The unchanged rule requires differences of at most 3 of 64 items in content preservation, uppercase events, and their intersection. This is operational matching, not a statistical equivalence test.

Base + uppercase instruction
61/64
Uppercase training · epoch 1
62/64
Uppercase training · epoch 2
64/64
Uppercase training · epoch 3
64/64
Normal training · epoch 1
0/64
Normal training · epoch 2
0/64
Normal training · epoch 3
0/64
All64 responses retained per condition. Base uses the uppercase instruction; trained models use the original case-preserving instruction. Counts come from development data already used for task feasibility. Final test items remain untouched.
Absolute difference from base + uppercase, in items (allowed: at most3 each)
CheckpointContentUppercaseBothGate
Epoch 1011Pass
Epoch 2213Pass
Epoch 3213Pass

The uppercase adapter was trained to produce capitals despite an explicit instruction to preserve case. That conflict is part of this task. A normal-case adapter, selected at the same epoch, controls some generic training drift. Neither matching nor a failure to match identifies a mechanism.

All checkpoints and cues, including exact case and lowercase compliance
Route · cueContentUppercaseBothExact normalExact uppercaseExact lowercase
base-0 · neutral61/640/640/6461/640/640/64
base-0 · uppercase62/6463/6461/640/6461/640/64
base-0 · lowercase60/640/640/640/640/6460/64
normal-1 · neutral63/640/640/6463/640/640/64
normal-1 · uppercase62/6463/6461/640/6461/640/64
normal-1 · lowercase60/640/640/640/640/6460/64
normal-2 · neutral64/640/640/6464/640/640/64
normal-2 · uppercase61/6463/6460/640/6460/640/64
normal-2 · lowercase61/640/640/640/640/6461/64
normal-3 · neutral64/640/640/6464/640/640/64
normal-3 · uppercase61/6463/6460/640/6460/640/64
normal-3 · lowercase61/640/640/640/640/6461/64
uppercase-1 · neutral62/6464/6462/640/6462/640/64
uppercase-1 · uppercase62/6464/6462/640/6462/640/64
uppercase-1 · lowercase59/640/640/640/640/6459/64
uppercase-2 · neutral64/6464/6464/640/6464/640/64
uppercase-2 · uppercase60/6464/6460/640/6460/640/64
uppercase-2 · lowercase61/640/640/640/640/6461/64
uppercase-3 · neutral64/6464/6464/640/6464/640/64
uppercase-3 · uppercase64/6464/6464/640/6464/640/64
uppercase-3 · lowercase62/640/640/640/640/6462/64

All counts, paired item results and provenance ↗ JSON

Reproducibility check: all192 newly generated base-model outputs exactly matched the earlier screen and confirmation. Geometry and steering calibration are complete; final-test evaluation remains pending.

Measured geometry · 64 direction items · one training seed

Do prompt and training changes point the same way?

The changes align more closely in later layers. At the answer boundary, cosine is -0.04 at layer 0, 0.62 at layer 14 and 0.86 at layer 27. The final-layer values on shared normal-case and uppercase continuations are 0.88 and 0.74. This does not establish that either direction controls capitalization; steering calibration is reported below; held-out evaluation remains pending.

Each point compares the mean activation change from an uppercase instruction (base + uppercase minus base + neutral) with the mean change from uppercase training (uppercase adapter minus normal-case adapter, both under the neutral instruction). A cosine of 1 means parallel directions, 0 orthogonal, and −1 opposite.

Showing: Last assistant-prefix token.

Last assistant-prefix token: prompt and training contrast cosine by layerAll 28 zero-based post-block layers. Cosine axis ranges from minus one to one. Points are cosine between mean prompt-minus-base and mean caps-minus-normal contrasts; vertical bars are paired item bootstrap 95 percent intervals. Exact values and undefined results appear in the table below.-1-0.500.5107142027Layer (zero-based, post-block)CosineLayer 0: -0.04233843889734848; 95% interval -0.05411096338343058 to -0.031544248562700106; 0 invalid bootstrap replicatesLayer 1: -0.04166032439703332; 95% interval -0.06846797816781204 to -0.01655395316606727; 0 invalid bootstrap replicatesLayer 2: -0.12190567359148874; 95% interval -0.13567407041460366 to -0.10725238264573467; 0 invalid bootstrap replicatesLayer 3: -0.19748041845560937; 95% interval -0.20775459510918895 to -0.1858577517332473; 0 invalid bootstrap replicatesLayer 4: -0.21845838130957426; 95% interval -0.22483071172193073 to -0.21162005814791934; 0 invalid bootstrap replicatesLayer 5: -0.22603662670124264; 95% interval -0.2307559967662495 to -0.22079248997824208; 0 invalid bootstrap replicatesLayer 6: -0.14219004590551101; 95% interval -0.14666768212471276 to -0.13697991213110208; 0 invalid bootstrap replicatesLayer 7: -0.08922496315576606; 95% interval -0.09475331219872132 to -0.08364223751461604; 0 invalid bootstrap replicatesLayer 8: -0.017721695352993404; 95% interval -0.022951894930409104 to -0.011969978433852875; 0 invalid bootstrap replicatesLayer 9: 0.0685232863476826; 95% interval 0.06620181931588343 to 0.07095268048469378; 0 invalid bootstrap replicatesLayer 10: 0.12253313307142259; 95% interval 0.11986402379432919 to 0.1253409211705846; 0 invalid bootstrap replicatesLayer 11: 0.16358163981454543; 95% interval 0.1608851672031625 to 0.16663563899095354; 0 invalid bootstrap replicatesLayer 12: 0.13147253538374407; 95% interval 0.12814165253717288 to 0.13495148958015926; 0 invalid bootstrap replicatesLayer 13: 0.2657855940358417; 95% interval 0.26281270871534285 to 0.26892087688406946; 0 invalid bootstrap replicatesLayer 14: 0.6230721610069516; 95% interval 0.6166836902505815 to 0.6287904672156956; 0 invalid bootstrap replicatesLayer 15: 0.6637609244600784; 95% interval 0.6577138815911552 to 0.6691933280820224; 0 invalid bootstrap replicatesLayer 16: 0.6237295948779112; 95% interval 0.6184219810505199 to 0.6285115694573243; 0 invalid bootstrap replicatesLayer 17: 0.656254016040036; 95% interval 0.6511794139237826 to 0.6605670569683947; 0 invalid bootstrap replicatesLayer 18: 0.6205199049329455; 95% interval 0.6162517768252056 to 0.6246030371094069; 0 invalid bootstrap replicatesLayer 19: 0.6909542283947511; 95% interval 0.6879391812951723 to 0.6937468296033126; 0 invalid bootstrap replicatesLayer 20: 0.6976847006455138; 95% interval 0.6941699034207046 to 0.7009775522197429; 0 invalid bootstrap replicatesLayer 21: 0.6866290794634042; 95% interval 0.6832170972884257 to 0.6897191402782843; 0 invalid bootstrap replicatesLayer 22: 0.7542529599113406; 95% interval 0.7498450425452393 to 0.7582492692115601; 0 invalid bootstrap replicatesLayer 23: 0.7293690172011119; 95% interval 0.725040852841573 to 0.733262367794976; 0 invalid bootstrap replicatesLayer 24: 0.7189140804260339; 95% interval 0.7147329865675842 to 0.7227254665896664; 0 invalid bootstrap replicatesLayer 25: 0.7809463009914543; 95% interval 0.774839483050126 to 0.7862417828300948; 0 invalid bootstrap replicatesLayer 26: 0.8202832258769468; 95% interval 0.8144848044497606 to 0.8256180046050577; 0 invalid bootstrap replicatesLayer 27: 0.8605953666456344; 95% interval 0.8544964542778067 to 0.8650439309517219; 0 invalid bootstrap replicates
Teal points and line: cosine between mean contrasts. Blue vertical bars: paired 95% percentile intervals from 2000 bootstrap samples of the same 64 items. All 28 layers are shown; intervals are pointwise, not simultaneous. Undefined values are omitted from the plot and retained in the numeric tables. A continuation state is averaged within each item before items are weighted equally; EOS is excluded.
These are exploratory geometric comparisons. Unequal cue content and token positions remain confounded. Forced shared continuations are controlled computations, not free generation. One training seed was used; intervals describe item variation only, not training-seed uncertainty. Geometry establishes no steering effect or causal mediation, and provides no final-test result.
Exact numeric values for all three measurements and all 28 layers
Last assistant-prefix token — all layers, original JSON precision
LayerCosine95% lower95% upperInvalid bootstrap replicates
0-0.04233843889734848-0.05411096338343058-0.0315442485627001060
1-0.04166032439703332-0.06846797816781204-0.016553953166067270
2-0.12190567359148874-0.13567407041460366-0.107252382645734670
3-0.19748041845560937-0.20775459510918895-0.18585775173324730
4-0.21845838130957426-0.22483071172193073-0.211620058147919340
5-0.22603662670124264-0.2307559967662495-0.220792489978242080
6-0.14219004590551101-0.14666768212471276-0.136979912131102080
7-0.08922496315576606-0.09475331219872132-0.083642237514616040
8-0.017721695352993404-0.022951894930409104-0.0119699784338528750
90.06852328634768260.066201819315883430.070952680484693780
100.122533133071422590.119864023794329190.12534092117058460
110.163581639814545430.16088516720316250.166635638990953540
120.131472535383744070.128141652537172880.134951489580159260
130.26578559403584170.262812708715342850.268920876884069460
140.62307216100695160.61668369025058150.62879046721569560
150.66376092446007840.65771388159115520.66919332808202240
160.62372959487791120.61842198105051990.62851156945732430
170.6562540160400360.65117941392378260.66056705696839470
180.62051990493294550.61625177682520560.62460303710940690
190.69095422839475110.68793918129517230.69374682960331260
200.69768470064551380.69416990342070460.70097755221974290
210.68662907946340420.68321709728842570.68971914027828430
220.75425295991134060.74984504254523930.75824926921156010
230.72936901720111190.7250408528415730.7332623677949760
240.71891408042603390.71473298656758420.72272546658966640
250.78094630099145430.7748394830501260.78624178283009480
260.82028322587694680.81448480444976060.82561800460505770
270.86059536664563440.85449645427780670.86504393095172190
Shared normal-case continuation — all layers, original JSON precision
LayerCosine95% lower95% upperInvalid bootstrap replicates
0-0.10024296895858946-0.13036891445879778-0.064866720844703190
1-0.09451565096415529-0.11197249330658145-0.074702944918210490
2-0.13117147464532333-0.14144488697452912-0.119728480965417970
3-0.08631444270952265-0.1033275052704698-0.069909111026885820
4-0.06349727340214792-0.07496171402052944-0.051385934187533160
50.214119085960270130.200908186256855750.225715300951541020
60.132763139537046130.11555421318283410.14937020854586760
70.262077146479460730.242795320731028440.28114151500874270
80.32369466592908240.30712546928680890.339583949000302640
90.26219071434201470.248673938673692250.275818467743305760
100.324132817204105530.31198667092352420.33599586577631120
110.3757627917209950.36286664554687730.38832495104783230
120.3774163337372650.365801776685015350.38887970168360930
130.44294918919204830.4347921320239040.45105990597052820
140.60471925445395010.59881490295272510.61002524325016070
150.66638919352700340.66141044950654330.67053786135198260
160.68810395396460330.68391801836917440.69153353942482230
170.71337827706047390.70963379348740430.71619124657579290
180.72788575673720450.72313371826358820.73189690725170580
190.78057746398289640.77711752828558480.78355054767612740
200.79284002287426370.78985685418210480.79543402070282180
210.78841380250472340.78520347619185860.79111533341083080
220.78298307736393680.77969365982137230.78580039513670710
230.78635771207364270.7819202989190410.79042284477449910
240.78902967720628810.7847189120748020.79314223826659710
250.81311943367967730.80768630495982070.81787927504079590
260.84424641733402130.83922446259935250.84839705221715970
270.87904335978925760.87139506488967840.88476984839752250
Shared uppercase continuation — all layers, original JSON precision
LayerCosine95% lower95% upperInvalid bootstrap replicates
00.022453832757696503-0.00146970770690499050.049691092096373240
1-0.002398468155298851-0.0222579787642797520.0189592349198864950
2-0.060035113175078245-0.08346476782301543-0.0395653518863402640
3-0.09317674303955209-0.12436663477000091-0.063668082514013330
4-0.06508759663295356-0.08657298381218052-0.044624105200728670
50.14525136241934760.115597273491294880.17212822132149490
60.118402263265295350.085873829482671530.151944813168043840
70.204558563762748140.17036860808933910.235292109744132870
80.247248374159397960.218775548412631440.272551638352317770
90.19071273684975680.174060425578252880.208426423988295940
100.287269680714759560.27075808779007370.30311690954297640
110.29536611907609260.27968583346929260.310348803839872230
120.302641875453888930.290206921134652850.315498406536325550
130.412446348375627770.40263456689876780.421794682580063640
140.49772345050200620.48968665627691580.50531212087495330
150.57426631860952950.56709951265469350.58126832452303980
160.5925167142746030.5850854418056310.59926117174438880
170.59981373192680.59293895798724010.60648498604582470
180.59165406686541520.58272990188502170.5997336482104780
190.63157583673563870.62340014917802110.63911874183786780
200.6587741661608390.65093278590823020.66640601306754410
210.66269984826579820.65483400143222310.66982305326891310
220.6655034586919180.65725700553302240.67214962814594880
230.70001517756119770.69320377495108380.70574998440517440
240.71709225684981990.70909551158168670.72324911668451770
250.72776627438414740.72031672032171180.73368844190150430
260.7335810444000040.72551735064192480.74025357867503060
270.7390955386028370.72647498538380920.74906818669027960

Intervals use finite bootstrap replicates; the invalid-replicate column reports omitted undefined replicates out of 2000. Undefined cosine means at least one mean contrast has zero norm.

Download geometry, bootstrap metadata and provenance ↗ JSON

Completed calibration · 19 September 2026 · 2,880 outputs

Removing capitals is not the same as restoring the answer.

All 40 nonzero candidates across four donor/recipient families have been evaluated. The frozen rule maximizes exact normal-case restoration subject to losing at most three content successes out of 64; families without improvement retain the zero intervention.

Selected interventions on reused calibration data
DirectionRecipientSelectionExact normalContent
PromptBase + uppercase promptLayer 20, scale -161/6461/64
PromptUppercase adapter + neutral promptNo calibrated intervention0/6462/64
TrainingBase + uppercase promptLayer 20, scale -161/6461/64
TrainingUppercase adapter + neutral promptNo calibrated intervention0/6462/64

Both selected interventions restore 61/64 exact normal answers in the prompted model, versus 0/64 without intervention, while content changes from 62/64 to 61/64. Every tested intervention into the trained model restores 0/64 exact normal answers. Yet the training direction at layer 20, scale −1 removes all uppercase events while retaining content on 62/64: outputs become mixed or title case. That does not meet the restoration objective.

All candidate counts, selections and provenance ↗ JSON

The first obstacle

Before comparing mechanisms,
make the behavior reliable.

Our small test asks a model to copy a supplied sentence, sometimes changing its letters to uppercase. The words should stay the same. That separates a style change from a change in content.

On these exploratory items, producing capital letters was easier than preserving the sentence. None of the three 1.5B configurations below passed the fixed feasibility screen. A later 7B screen passed, and then passed the frozen 48-item confirmation. There is no substantive prompt-versus-training mechanism result yet.

Measured · Qwen2.5-1.5B-Instruct

All caps can hide the wrong words.

Figure data ↗ JSON

With mixed-case input, the uppercase instruction produced the requested broad style on 16 of 16 items, but preserved the content on only 12 of 16. Supplying uppercase input improved uppercase copying, while neutral copying lost content.

Content preservedStyle eventBoth on the same response

copy-v2 · gate failed

Mixed-case input

Repetition penalty 1.1

Neutral

Content
16/16
Style
16/16
Joint
16/16

Uppercase

Content
12/16
Style
16/16
Joint
12/16

Lowercase

Content
16/16
Style
16/16
Joint
16/16

copy-v3 · gate failed

Uppercase input

Repetition penalty 1.1

Neutral

Content
13/16
Style
14/16
Joint
11/16

Uppercase

Content
15/16
Style
16/16
Joint
15/16

Lowercase

Content
16/16
Style
16/16
Joint
16/16

copy-v4 · gate failed

Mixed-case input

Repetition penalty 1.0

Neutral

Content
16/16
Style
16/16
Joint
16/16

Uppercase

Content
12/16
Style
16/16
Joint
12/16

Lowercase

Content
16/16
Style
16/16
Joint
16/16
Each bar counts successes out of 16 responses. Every response remains in its denominator. These are the same 16 exploratory items across all three versions, reused during adaptive diagnosis—not three independent samples. No uncertainty intervals or held-out generalisation estimate are claimed.
Exact counts, stricter scoring, and the pass rule
All conditions; no responses discarded
Version · instructionContentStyleJointExact requested case
copy-v2 · neutral16/1616/1616/1615/16
copy-v2 · uppercase12/1616/1612/1612/16
copy-v2 · lowercase16/1616/1616/1616/16
copy-v3 · neutral13/1614/1611/1611/16
copy-v3 · uppercase15/1616/1615/1615/16
copy-v3 · lowercase16/1616/1616/1616/16
copy-v4 · neutral16/1616/1616/1615/16
copy-v4 · uppercase12/1616/1612/1612/16
copy-v4 · lowercase16/1616/1616/1616/16

Content compares the full output to the reference after lowercasing and trimming outer whitespace. Words, internal spaces, and punctuation must match. Joint requires both content and the style event. Exact requested case additionally requires the exact target text, including case.

The screen required content preservation on at least 14/16 items in each condition, uppercase and lowercase style on at least 14/16 each, and no more than 2/16 neutral treatment events. These are separate marginal gates, not a joint-success gate or statistical equivalence test. In v3, neutral content was 13/16, neutral style 14/16, and their intersection only 11/16.

One actual v2/v4 error · calibration-001

Input: Spoon comes after leaf in alphabetical order.

SPONGE COMES AFTER LEAF IN ALPHABETICAL ORDER.

The uppercase style succeeds. Copying fails: “spoon” became “sponge.”

A decoding check did not resolve this bottleneck: changing the repetition penalty from 1.1 to 1.0 changed none of 48 paired greedy outputs in v4; its control outputs also matched the historical v2 outputs. This is a text comparison, not evidence that the logits were identical.

New checkpoint · preliminary 7B screen

A larger model clears the first screen.

Qwen2.5-7B-Instruct, using the v2 mixed-case copying setup, passed the same marginal gates on the same 16 exploratory items. All three style counts were 16/16; content and joint success were lower, as shown below.

Preliminary 7B screen · 16 responses per instruction
InstructionContentStyleJointExact requested case
Neutral14/1616/1614/1614/16
Uppercase15/1616/1615/1615/16
Lowercase14/1616/1614/1614/16

This is a descriptive comparison between model checkpoints, not a causal estimate of the effect of parameter count. The reused screen alone is insufficient: a separate 48-item confirmation also passed (below). No mechanism conclusion follows from passing a copying gate.

7B screen data & provenance ↗ JSON

Measured · 48 previously unused calibration items

The fixed protocol passes confirmation.

The same 7B model, instructions and decoding settings were applied without changes to 48 additional items. Each condition exceeded the frozen 42/48 content threshold. Uppercase style succeeded on 47/48 outputs, lowercase style on 48/48, and neutral uppercase events were 0/48.

7B confirmation · all responses retained
InstructionContentStyleJointExact requested case
Neutral47/4848/4847/4847/48
Uppercase47/4847/4846/4846/48
Lowercase46/4848/4846/4846/48

Uppercase content and style each succeeded on 47 items, but their intersection was 46: they failed on different responses. The gate establishes feasibility for this templated copying task, not prompt/training equivalence, error-free copying or general reasoning. These items were unused before confirmation; they now belong to calibration, not the final test.

Confirmation data & provenance ↗ JSON

Planned · conditional on feasibility

Compare the routes, then test transfer.

The core experiment needs two ways to elicit comparable behavior. A stronger uppercase rate alone would make the comparison ambiguous.

A / instruction

Change the prompt

Keep the model’s weights fixed. Compare a neutral instruction with an uppercase instruction on the same items.

B / training

Change a small set of weights

Train an uppercase adapter and a normal-case control adapter on matched content. Evaluate both with a neutral instruction.

  1. Match observable performance. Use calibration data separate from training and final testing to find conditions with similar style and content success. A mismatch must remain visible.
  2. Measure internal changes. Compare changes in the model’s residual activations—the running internal representation—at each layer. Examine the answer boundary and identical supplied answer tokens separately.
  3. Try the change in the other route. Add a prompt-derived activation direction to the trained model, and a training-derived direction to the base model. Include opposite signs, random directions and utility checks.
  4. Evaluate once on held-out items. Freeze analysis choices first. Similar directions are not proof of a shared mechanism; functional transfer is a further test, with its own alternative explanations.

What did not pan out

A smaller task still needed scrutiny.

Earlier 0.5B and 1.5B diagnostics failed a supplied-word alphabetical task and case-copying screens. Copy-specific instructions improved copying but did not pass the gates; reversing input case shifted the failure, and disabling the repetition penalty did not fix it.

A technical smoke test captured activations and performed one optimizer step. That establishes limited plumbing, not a completed training comparison.

What this can tell us

A measurement lesson, so far.

Measure content and style on the same outputs before treating style success as competence. These results diagnose this small, templated test; they do not establish how prompting differs mechanistically from training.

Capitalization is a cheap test behavior, not an alignment, personality, consciousness or welfare endpoint.

Context & provenance

This builds on existing work.

Persona Vectors already relates training-induced activation changes to prompt-derived directions and uses those directions to steer trained models. Concept Ablation Fine-Tuning (CAFT) already extracts training-related activation differences on shared text and tests inference ablations. Geometry and one-way transfer are established ideas, not a novelty claim here.

The narrower proposed question is when similarity between behavior-matched routes agrees—or disagrees—with transfer in both directions. The current geometry describes directional similarity; the suppression calibration below tests both directions; held-out functional transfer remains pending.

The figure dataset includes metric definitions, all counts, model revision, run identifiers, original workspace paths and SHA-256 hashes for source reports and code. Workspace paths in that file are provenance identifiers, not links in this standalone site. The build is reproducible from the workspace with python3 site/build.py.