StructuraRESEARCH WORKBENCHMID-SEMESTER REVIEW / SEPTEMBER 2026
Structure-aware preference alignment

Beyond preference.
Inspect the constraints.

Flow-matching preference optimization with decomposed semantic and structural evidence.

Exploratory automated-signal ablations — not human-calibrated
128Real preference pairs
5 / 5Completed training runs
640Recorded optimizer steps
4.127 GiBPeak allocated T4 memory

The research question

Implemented & pilot-trained

Can component-level constraints improve prompt and structural correctness over ordinary human-preference optimization?

01 / EVIDENCE

Preference pairs

One prompt, a human-preferred image and a rejected image. Preserve identities and hashes.

02 / DECOMPOSITION

Eight signals

Semantic, presence, color, attribute, count, relation, geometry and visual-quality proxies.

03 / OPTIMIZATION

Frozen reference + LoRA

SD3 Medium flow backbone. A shared pooled OT schedule and human-gradient conflict protection.

04 / EXPERIMENT

Controlled ablations

Five objectives, identical starting adapters and 128 steps, followed by 360 matched generated outputs.

What is different in our design?

Instead of collapsing every signal into a single preference score, preserve separate signed constraint margins, applicability masks and weights.

  • Ignore constraints that do not apply to the prompt.
  • Use a minimum-margin deadband to suppress tiny score differences.
  • Limit conflicting auxiliary influence so it cannot reverse the human preference direction.
  • Keep reliability gates separate from conflict protection.

These constraints supervise a shared flow-preference logit; they are not independent detection heads or direct pixel-space losses. OT is the coupling used within training, not a separately trained model.

What this pilot establishes

  • All five real-model training runs completed on a T4.
  • Same data, initialization, schedule and base-model revision across runs.
  • All 360 matched evaluation images and decomposed scores were produced.
It does not establish superiority over the base papers or a validated human-calibrated reward system. Most automatic metric differences remain uncertain.

Reliability weights are one for active signals in this campaign. The final reliability-gated method remains a later experiment.