DesignJournal

How Agent-Driven Creative Trace Mining Reshapes Design Behavior in a Deployment Study of Intent-Aligned Multimodal Design

Seung Won Lee1,*, Jiin Choi1,*, Yejin Yun1,*, Semin Jin1, Yugyeong Jang1, Geunmin Hwang2, Sang Woon Park2, Seonghoon Ban2, Kyung Hoon Hyun1,†
1Design AI Lab, Department of Interior Architecture Design, Hanyang University, Seoul, Republic of Korea
2RECON Labs Inc., Republic of Korea
*Equal contribution  ·  Corresponding author  ·  International Journal of Human–Computer Interaction
Paper soon Code

TL;DR

Agents don’t just add actions to design work — they reconfigure the sequence of it.

Designers capture inspirations constantly — sketches, photographs, voice memos, notes — but these creative traces rarely flow back into creation. DesignJournal consolidates multimodal artifacts into one workspace and mines them with an agentic pipeline: structured grouping by three role-differentiated agents (ECHO, INSIGHT, SPARK), belief-guided prompt refinement, and iterative intent updates. To measure what this does to real practice, we ran a 10-day in-the-wild deployment with eight professional designers and built a framework that classifies 5,838 design actions into a dual-origin taxonomy. Agent support significantly reduced spatial repositioning — concentrated in small moves — and shifted transition structures from long repositioning chains toward relational encoding sequences.

DesignJournal system overview: the interface consolidates multimodal design artifacts — text, sketches, images, and audio — into a timeline and canvas with agent-based grouping and generation.
Figure 1. DesignJournal system overview — the interface consolidates multimodal design artifacts (text, sketches, images, and audio) into a chronologically revisitable timeline and a spatially organizable canvas, and turns them into intent-aligned generation.
5,838

classified design actions from 14,075 raw events

10 days

in-the-wild deployment, 5 days per condition

8

professional designers, within-subject

−43%

Relocate actions with agent support (p = .039)

The problem

Records accumulate. The pipeline from capture to reuse breaks.

Designers work across seven to eight parallel tools. Ideas scatter across apps, boards, and folders, and a findability bottleneck recurs every time reuse is attempted. Four obstacles motivate four design goals.

DG1

Consolidate scattered traces

Traces scatter across parallel tools and become a findability bottleneck during goal-directed reuse. A single repository must keep multimodal traces revisitable both chronologically and spatially.

DG2

Preserve explicit relations

Pinboards either keep connections off-canvas or lack model-driven interpretation, so relational context erodes as the collection grows. Relationships between items must be explicit and persistent at scale.

DG3

Turn records into evidence

A single text prompt rarely conveys implicit intent; text-centric input narrows interpretive range and induces fixation. Generation should be grounded in accumulated, time- and relation-structured records — not isolated prompts.

DG4

Align intent across stages

Existing multiturn T2I offers no group-level reuse, and generative output carries documented fixation risk. Support needs distinct roles for consolidation, reinterpretation, and divergence — each keeping intent aligned across iterations.

RQ1

Which changes in design action patterns emerge when agent-based affordances are introduced?

RQ2

Which new human–AI collaborative loops arise, and how do they replace or complement existing practices?

RQ3

How do designers perceive appropriate agent-mediated workflows in terms of authorship, traceability, and agency?

The system

Four parts, from capture to intent-aligned generation.

Records enter as notes, images, drawings, tables, to-do lists, or voice memos. They are stored as structured JSON, arranged on a canvas with a chronological timeline alongside, and composed into prompts by the agents.

part 01 · DG1, DG3

Goal & Record Input

A Global Goal plus up to three Local Intents, alongside unified multimodal capture — notes, images, audio, sketches — each annotated with Captions, Reactions, and Comments stating why it was saved.

part 02 · DG1, DG2

Canvas Organization

The left timeline tracks items chronologically; the central canvas lets designers place, link, and cluster records at the group level so relations persist as the collection grows.

part 03 · DG4

Agent Processing

Three agents infer implicit intent by computing similarities over extracted descriptions, then generate and refine targeted questions. Each item carries a high/low belief signal.

part 04 · DG3, DG4

Generation & Evaluation

Group-level prompts produce alternatives that the designer compares, selects, and saves back as new inspiration — with the Goal and Local Intents refined after each feedback loop.

The DesignJournal pipeline, from goal and record input through canvas organization and agent processing to generation and evaluation.
Figure 2. Pipeline of DesignJournal.
The DesignJournal interface: set goal, agent group mode, generate idea, and edit and generate.
Figure 3. (a) Set Goal — declare the Global Goal and add Local Intents, with system-generated suggestions from the current design context · (b) Agent (Group Mode) — records organized into ECHO, INSIGHT, and SPARK groups · (c) Generate Idea — color-highlighted prompts with targeted refinement questions · (d) Edit & Generate — localized, point-guided edits that preserve composition, lighting, and scene consistency.
Artifact types in DesignJournal with caption, reaction, and comment functions.
Figure 4. Artifact types and the (a) Caption, (b) Reaction, and (c) Comment functions — the annotations that record why an item was kept.

Three agents

Roles defined by how a record relates to declared intent.

The same record can land in a different group depending on the current Goal and Local Intent — the assignment is relational, not categorical. Each agent synthesizes its own prompt from its group, so the designer works with three parallel alternatives from three distinct perspectives.

Pick an agent

ECHO

Consolidating decision-ready knowledge

Records it pulls in

    Synthesized prompt → refinement question

    Refinement question

    Illustrative example from the cafe lounge brief. Prompts are strictly grounded in the collected descriptions — exaggeration is prohibited — and questions target only visually verifiable attributes (color, material, form, layout, composition, style).

    Group-based generation and iterative refinement workflow with agents: trigger generation, view images, answer suggested questions, re-generate, save, and see refined goals.
    Figure 5. Group-based generation and iterative refinement. (a) Trigger generation per group · (b) Synthesized prompt + images · (c) Suggested Questions aligned to goals and intents · (d) Re-Generate on the answer · (e) Add to My DesignJournal · (f) Goal and Local Intents refined in the Set Goal panel.

    Belief-guided prompt refinement, batched.

    Multiturn back-and-forth burdens cognition, so the refinement agent uses a batch-questioning scheme: it proposes questions about multiple attributes across one or more entities at once, collects a single response, and updates the prompt — preserving the benefit of multiturn interaction without the latency.

    Artifacts are split by belief. High-belief artifacts drive questions that encourage diversity around well-captured elements; low-belief artifacts target ambiguities to resolve underspecified parts of the prompt. The agent then applies the response while minimizing edits — preserving core context and tone, updating only the changed attributes.

    Prevalidation · IDEA-Bench T2I

    w/o agent (baseline) 75.00
    w/ agent (ours) 83.33

    With the Prompt Refinement Agent, generation reaches the IDEA-Bench T2I state of the art (Flux-1, 83.33) under MLLM-based automated evaluation.

    Analysis framework

    A dual-origin action taxonomy over 5,838 actions.

    To ask whether agents merely add actions or reconfigure sequences, we needed a classification that separates designer-initiated manipulation from system-mediated generation. 14,075 raw events → 7,293 clean events → 5,838 mappable actions.

    State-based inference

    7 categories

    Designer-initiated actions on seven artifact types, classified by comparing each instance’s current state St against St−1 across seven deltas — life cycle, Δchar, Δpos, attributes, annotations, relation endpoints, and hierarchy.

    Create Elaborate Relocate Relate Structure Prune Interact

    Explicit system events

    4 categories

    AI-mediated actions detected from system-level markers, capturing four distinct AI roles across the generation and intent-alignment loop.

    AgentGen PromptGen ImageEdit IntentEdit

    Hover any category for its definition.

    method 01

    Action count comparison

    Wilcoxon signed-rank tests on paired raw counts per category, for volume shifts between conditions.

    method 02

    n-gram sequence analysis

    Patterns at n = 2–5 per participant, with concurrent action pairs expanded into all 2k permutations and equally weighted.

    method 03

    Transition divergence (JSD)

    First-order Markov transition matrices per condition, compared by Jensen–Shannon divergence with a 5,000-iteration permutation test.

    method 04

    Collaborative loops

    Recurring sequences of alternating designer-initiated and agent-mediated actions, characterized for RQ2.

    Deployment

    Ten days, in the wild, each designer their own baseline.

    Laboratory tasks can measure short-term performance, but they cannot reveal whether records accumulate into reusable creative traces, whether intent alignment holds over time, or whether designers actually change how they revisit and recombine past artifacts.

    Condition A

    DesignJournalbaseline

    The unified multimodal workspace with manual organization — timeline, canvas, relational linking, and direct prompt-based generation.

    Condition B

    DesignJournal+Agent

    The same workspace plus the agentic pipeline: automatic ECHO/INSIGHT/SPARK grouping, belief-guided prompt refinement, and iterative intent updates.

    Participants

    Eight professional designers (1–8 years’ experience, 30–60 h/week) from design agencies and in-house corporate teams — industrial, product, furniture, and home furnishings design. Within-subject, so each serves as their own baseline. IRB-approved (HYU-2025-138).

    Task

    Two standardized client-style briefs — a specialty cafe lounge (space) and a kitchenware collection (product) — one per condition, counterbalanced. Deliver a final outcome plus a moodboard with written rationale suitable for client communication.

    Protocol

    Five working days per condition, each followed by a post-condition survey and semi-structured interview. Designers kept their familiar tools (CAD, Adobe, whiteboards) but funneled every inspiration through the study tool for traceability, with lightweight daily documentation.

    Experimental timeline of the 10-day deployment study.
    Figure 7. Experimental timeline of the 10-day deployment study.
    Experimental briefs: a space brief for a specialty cafe lounge and a product brief for a kitchenware collection.
    Figure 8. The briefs: (a) a specialty cafe lounge and (b) a kitchenware collection.

    Results · RQ1

    Less shuffling, differently sequenced.

    Total action volume did not differ significantly between conditions (W = 6, p = .109). What changed was which actions and, more tellingly, in what order.

    Action distribution

    Relocate fell from M = 197.5 to M = 112.1.

    W = 3, p = .039, r = 0.83 — a large effect, with six of eight designers decreasing. No other shared category changed significantly.

    Magnitude

    The drop concentrated in small displacements.

    Small: 75.5 → 27.6 (p = .047, d = 0.81), significant in 100% of 1,000 Monte Carlo threshold perturbations. Medium and large did not reach significance.

    Sequential structure

    Transition structure differed significantly.

    Group-level permutation test: observed M = 0.289 vs. null M = 0.248, Z = 2.24, p = .010. Individual-level corrected JSD: p = .031, r = +1.00.

    Where the repositioning went

    Relocate stratified into three magnitude levels at the 33rd and 67th percentiles of the pooled displacement distribution.

    Magnitude baseline M (SD) +Agent M (SD) p d MC robustness
    Small 75.5 (58.3) 27.6 (15.3) .047* 0.81 100% (±15 pct.)
    Medium 64.4 (51.6) 39.0 (22.0) .109 0.54 4.9%
    Large 57.6 (36.2) 45.5 (23.7) .461 0.37 0.0%

    Wilcoxon signed-rank tests. MC robustness = proportion of Monte Carlo perturbation runs preserving significance under threshold variation. The same stratification applied to Elaborate yielded no significant differences at any magnitude.

    Interviews explained the mechanism. P6: in the baseline, “I had to manually move images here and there, but the agent grouped them automatically.” P4 described the baseline as requiring “planning even the grouping part from scratch on a blank slate,” while agent-mediated grouping “took care of more than half of that organizational work.” P1: “grouping through ECHO saved time, especially when making final decisions and organizing.”

    Repositioning chains attenuate as n grows

    Pure Relocate chains dominate the baseline at every order. With agent support they shrink at each one — and the gap widens the longer the chain.

    Change n
    Chain length

    Relocate → Relocate

    DesignJournalbaseline 29.92%
    DesignJournal+Agent 18.40%
    Difference −11.5 pp the single largest shift in the transition profile

    Relocate chain attenuation

    In the baseline, Relocate → Relocate dominated all 2-grams and stayed high through higher orders. Under agent support every order dropped. The baseline top-10 also contained Relocate → Elaborate (1.84%) and Elaborate → Relocate (1.76%) — workflows where spatial arrangement preceded content elaboration. Both were reduced.

    Relational patterns interleave

    Relate counts did not differ (p = .688), but their position did. Create → Relate entered the +Agent top-10 at 2.98%; Relate → Relate rose to 6.38% and Relate → Relocate to 5.82%. At higher orders, pure Relate chains did not grow — instead interleaved Relate–Relocate 3-grams displaced Relocate-dominated ones. Relational work became interwoven with spatial activity rather than occurring in isolated bursts.

    Participant-level action sequences in DesignJournal +Agent, one row per participant.
    Figure 11. Participant-level action sequences in DesignJournal+Agent. Each row is one participant (P1–P8); the x-axis is the action index.
    Participant-level action sequences in DesignJournal baseline, one row per participant.
    Figure 12. The same participants in DesignJournalbaseline — note the comparatively dense Relocate bands (gray).

    Results · RQ2

    New collaborative loops — and a new place to start.

    AgentGen actions were infrequent (5.84% of all actions), yet they recurred at the core of the generation–refinement loops. Their influence is structural, not volumetric.

    Entry point shift

    Baseline workflows began with the designer’s own manual initiation — keyword derivation, moodboard assembly, or sketching. With agent support the loop started earlier, from the collected material itself, replacing those manual routines.

    “I started from references rather than sketches from the beginning because the agent could work with collected materials.”

    P1 · usually starts by sketching

    Without the agent, “I had to derive keywords and start categorizing by myself from scratch.”

    P2 · on the manual entry point

    “Even when I didn’t have a concrete image in mind, I could use the Generate Idea function right away.”

    P5 · replacing Pinterest moodboarding

    Divergent loops

    expand before commitment

    Visible in the transition data: AgentGen → AgentGen entered the +Agent top-10 2-grams at 2.30%. Among the participants who used it, this reflected sequential generation (M = 34.10%, SD = 22.33%) — producing multiple alternatives before selecting.

    In P2’s progression, the agent first acted as a divergent catalyst, generating kitchenware alternatives through sequential SPARK triggers.

    Convergent loops

    narrow through exchange

    Iterative cycling through intent articulation, agent generation, and image editing until a cohesive design settles. P2 refined and committed selected outputs over five days, mobilizing INSIGHT and a final round of SPARK.

    P1 used ECHO and INSIGHT for direction validation, progressively developing standing tables, evening mood lighting, rounded furniture, and cocoon seating — converging on a cozy cafe with natural materials and terracotta floor accents.

    Final design outcomes and corresponding records for P1 and P8 using DesignJournal +Agent.
    Figure 9. Final outcomes and their records for (a) P1 and (b) P8 under DesignJournal+Agent — P8 progressed from concept definition to detailed design through agent-assisted editing, arriving at mezzanine seating, soft rounded lighting, and clean-lined bar tables.
    Design progression of P2 in DesignJournal +Agent over five days and 554 actions.
    Figure 13. P2’s progression in DesignJournal+Agent (product brief, 554 actions over 5 days) — divergent SPARK triggers early, convergent refinement later.

    Results · RQ3

    Awareness of the agent reached back into recording.

    Knowing an agent would read the journal changed what designers were willing to put in it — an upstream effect on capture, not just on generation.

    In the baseline, “I didn’t have to worry about the agent, so I had more freedom in uploading records.” With the agent, “I felt constrained because I was thinking about what the agent would see.”

    P5 · on self-censored capture

    “If I could distinguish which elements the agent is looking at, I would have used it better — some memos I wouldn’t want it to read, others I would.”

    P2 · on selective visibility

    The takeaway isn’t that agents make designers faster — total action volume barely moved. It’s that agents redistribute design effort. Small-scale spatial fiddling, the manual proxy for organizing thought, gets absorbed by automatic grouping; what remains reorganizes around relational encoding and generation–refinement exchange. Agent support also extended designers’ creative reach beyond independent exploration — while raising open questions about authorship, traceability, and how much of the journal a designer actually wants read.

    Deployment architecture

    Every action, timestamped and traceable.

    DesignJournal captures and processes all designer activity through a serverless cloud stack — the source of the 14,075 raw interaction events behind the analysis.

    Serverless cloud stack for the deployment experiment.
    Figure 6. Serverless cloud stack for the deployment experiment.

    BibTeX

    @article{lee2026designjournal,
      title   = {How Agent-Driven Creative Trace Mining Reshapes Design Behavior
                 in a Deployment Study of Intent-Aligned Multimodal Design},
      author  = {Lee, Seung Won and Choi, Jiin and Yun, Yejin and Jin, Semin
                 and Jang, Yugyeong and Hwang, Geunmin and Park, Sang Woon
                 and Ban, Seonghoon and Hyun, Kyung Hoon},
      journal = {International Journal of Human--Computer Interaction},
      year    = {2026},
      note    = {Under review},
    }

    Under review at IJHCI. Volume, pages, and DOI will be filled in once the paper is accepted and the official citation is released.