Preference pairs
One prompt, a human-preferred image and a rejected image. Preserve identities and hashes.
Flow-matching preference optimization with decomposed semantic and structural evidence.
Exploratory automated-signal ablations — not human-calibratedCan component-level constraints improve prompt and structural correctness over ordinary human-preference optimization?
One prompt, a human-preferred image and a rejected image. Preserve identities and hashes.
Semantic, presence, color, attribute, count, relation, geometry and visual-quality proxies.
SD3 Medium flow backbone. A shared pooled OT schedule and human-gradient conflict protection.
Five objectives, identical starting adapters and 128 steps, followed by 360 matched generated outputs.
Instead of collapsing every signal into a single preference score, preserve separate signed constraint margins, applicability masks and weights.
These constraints supervise a shared flow-preference logit; they are not independent detection heads or direct pixel-space losses. OT is the coupling used within training, not a separately trained model.
Reliability weights are one for active signals in this campaign. The final reliability-gated method remains a later experiment.
All 128 frozen training pairs are available here. These are dataset images, not outputs from the new adapters.
| Signal | Preferred | Rejected | Delta | Eligible |
|---|
Delta = preferred − rejected. A negative delta indicates evaluator disagreement with the human ranking, not a proven labeling error. Eligible means applicable and |delta| ≥ 0.02, before conflict protection.
This is a replay of actual saved metrics, not a currently running training job.
Steps with a nonzero logged contribution, out of 128.
| Base | SD3 Medium, pinned revision; frozen reference behavior | Policy | LoRA rank 4, FP32 trainable weights |
|---|---|---|---|
| Schedule | 128 steps; checkpoint every 16 | Data / seed | Same 128 pairs / seed 42 |
| Learning rate | 0.00005 | Memory | FP16 base; 256 px; batch 1 |
| OT | Pool 8, identical assignment hash | Initialization | Restored before each run |
Six outputs use the same prompt, seed, resolution and inference schedule. These are actual SD3 Medium and trained-LoRA outputs from the completed T4 evaluation.
This hosted copy displays recorded local build validation. Rechecking the research files requires the local Python server on the project computer.
Recorded local build validation passed. This hosted copy does not perform a fresh local-file integrity check. No credentials or human-calibration key are included.
60 prompts × 6 models = 360 images
This supports technical feasibility and identifies a signal worth follow-up. It is not a claim of superiority or a human-calibrated result.