01 / One-Sentence Positioning

项目首页与核心概述

Resonant Scenes

I designed an interactive experience that turns private emotions into something that can be seen, triggered and heard.

Placeholder image for collective emotion interface
04

Iteration 01: Camera-Based Gesture Recognition and Particle Feedback

The project began with camera-based gesture recognition. Hand movement, opening and waving were mapped to particle displacement, expansion, aggregation and colour changes in the browser. This stage confirmed an important hypothesis: users were willing to participate in emotional expression through bodily movement rather than through buttons alone. However, the procedural two-dimensional visuals were primarily composed of particles, lines, ripples and geometric forms. Even when emotions were assigned different colours, speeds and motion rules, the result often still resembled a technical demonstration. I therefore realised that an emotional experience requires not only dynamic feedback, but also a visual narrative capable of holding events, characters and spatial relationships.

Iteration 01: Camera-Based Gesture Recognition and Particle Feedback — image 1

Shame v1

Iteration 01: Camera-Based Gesture Recognition and Particle Feedback — image 2

Sadness v1

Iteration 01: Camera-Based Gesture Recognition and Particle Feedback — image 3

Excitement v1

05

Iteration 02: MPU6886, CoreS3 and the Interactive Sphere

To move the interaction beyond an on-screen gesture and into a physically held experience, I designed a shakeable spherical object. The prototype used an MPU6886 motion sensor to read acceleration, rotation and shaking behaviour. A CoreS3 board and its illuminated display were placed inside a transparent or translucent plastic sphere. Reflective fabric or similar reflective material was intended to diffuse and multiply the screen light within the sphere, creating a sense of enclosure and internal glow. Shaking the sphere could drive two stages of the experience: before the AI image was generated, it controlled browser-based emotional particles and procedural two-dimensional animations; after the AI image appeared, it continued to control image movement, colour, local effects and subsequent visual changes. This prototype strengthened the bodily relationship between holding, weight, inertia and shaking. It also introduced the symbolic idea of “holding an emotion in one’s hands”. However, powering the CoreS3 board, the MPU6886 and continuous illumination required a battery. The battery, board, wiring, mounting structure and plastic shell collectively increased the size of the sphere. An object intended to be held naturally gradually became oversized, heavy and dependent on maintenance. The hardware also introduced new barriers: users had to be physically present with the device; the system required charging and maintenance; different grip sizes were constrained by the sphere; the experience could not be shared remotely with ease; every demonstration depended on complete hardware availability. This failure did not make me abandon embodied interaction. Instead, it clarified what the project actually needed to preserve: not the sphere itself, but the relationship in which shaking causes an emotional scene to respond to the body. The next design goal therefore became: Remove the dedicated sphere while preserving the shake; remove the device barrier while preserving bodily participation. The interaction was ultimately transferred to the motion sensors already embedded in an ordinary smartphone. A user now only needs to open a URL, grant motion permission and gently shake the phone to enter the same visual, sonic and five-stage emotional experience.

Iteration 02: MPU6886, CoreS3 and the Interactive Sphere — image 1

Synchronize hardware and the web interface using Wi-Fi communication.

Iteration 02: MPU6886, CoreS3 and the Interactive Sphere — image 2

Interaction hardware architecture diagram for Plan 2: using RGB LED strips to emit light in a color corresponding to each detected emotion.

06

When Complexity Became Visible: From Local Generation To Cloud Intelligence

Alongside the hardware experiments, I also explored a locally deployed ComfyUI workflow for image generation. Local deployment allowed control over model nodes, workflow structure and certain visual parameters, while reducing direct dependence on external services. However, the project did not simply require the generation of a generic “emotional image”. The system needed to understand what had specifically happened to the participant, how the event related to the emotion, what spatial metaphor was appropriate, and which actions or objects should appear in the scene. For example, “I received the opportunity I had been waiting for, yet suddenly felt empty” and “I lost a relationship and therefore felt empty” may both be classified as numbness or loss, but they require entirely different narrative images. A local image workflow alone could not consistently perform this event-level semantic translation. The final pipeline therefore became: Specific personal event → language-model interpretation of event–emotion relationships → structured visual prompt construction → Seedream cloud image API → personalised emotional scene Cloud generation improved the system’s ability to understand concrete events, interpersonal relationships and visual metaphors. It also introduced network dependence and a waiting period that could approach thirty seconds. This technical trade-off directly led to the final two-layer visual architecture: an immediately available library of fourteen pre-generated emotional scenes; followed by a personalised AI image related to the user’s event.

When Complexity Became Visible: From Local Generation To Cloud Intelligence — image 1

Call the APIs for the semantic model and the image generation model within the code snippet.

When Complexity Became Visible: From Local Generation To Cloud Intelligence — image 2

Image1 created using ComfyUI based on users' experiences: due to its limited semantic understanding, it can only generate abstract images constrained by the prompts I designed.

When Complexity Became Visible: From Local Generation To Cloud Intelligence — image 3

Image2 created using ComfyUI based on users' experiences: even though it generates images faster than cloud-based solutions, it fails to achieve the desired therapeutic effect.

When Complexity Became Visible: From Local Generation To Cloud Intelligence — image 4

When I first started API to generate image, I constrained the output direction for each emotion; this resulted in nearly identical images for different experiences sharing the same emotional label, making it impossible to achieve the desired interactive effect.

When Complexity Became Visible: From Local Generation To Cloud Intelligence — image 5

So I had the semantic understanding model send specific image-based narratives to the image generation model based on the user's experiences; however, common issues associated with AI image generation still arose during the debugging phase.

When Complexity Became Visible: From Local Generation To Cloud Intelligence — image 6

Number of API calls made during debugging: DeepSeek is a highly cost-effective semantic understanding model.

07

Before AI Image Appeared: From Procedural Animation to Image-Led Emotional Scenes

The waiting experience before the personalised AI image appeared underwent two major transformations. The first stage relied entirely on browser-generated two-dimensional animation, including particles, curves, ripples, colour fields and geometric motion. These effects responded immediately to movement, yet lacked sufficient visual refinement and narrative depth. In the second stage, I created a library of fourteen emotional scenes through a unified prompt system. Browser code was then used to animate local regions of each image or modify specific visual properties of the image itself. This approach retained the strengths of both systems: AI-generated imagery provided characters, spaces, objects and narrative relationships; browser-based code provided immediate, repeatable interaction without requiring an additional API call. The waiting screen therefore ceased to be a temporary placeholder and became a complete affective prelude system.

Before AI Image Appeared: From Procedural Animation to Image-Led Emotional Scenes — image 1

Browser-generated image 1:Excitement

Before AI Image Appeared: From Procedural Animation to Image-Led Emotional Scenes — image 2

Browser-generated image 2:Shame

Before AI Image Appeared: From Procedural Animation to Image-Led Emotional Scenes — image 3

Final version: 14 emotional scenes through a unified prompt system Sharing a consistent visual language: warm beige, desaturated blue-grey and restrained warm illumination; paper-like, watercolour or subtly grainy surface textures; quiet cinematic spaces with clear foreground, middle-ground and background relationships; a consistent minimal stick-figure character; everyday objects used as emotional metaphors; restrained negative space without visual emptiness; emotion communicated through distance, lighting and object relationships rather than exaggerated facial expression. The stick figure functions as a neutral emotional vessel. It has no specific identity, age or facial expression, allowing different participants to project themselves into the scene.

10

Sound Design: Ambient Bed, Trigger Sound and Emotional Progression

The sound design evolved from a barely perceptible synthesiser tone played on each shake into a complete emotional sound system. The final structure consists of three layers: Ambient bed: a continuous synthesiser layer, white noise, rain texture or atmospheric room tone; Action transient: a clearly perceptible impact, droplet, crack, low-frequency hit or static pulse produced by a shake; Emotional tail: chords, bells, reverberation and delay that reconnect the event sound to the ambience. During each valid shake, the ambient bed is briefly reduced to create auditory space for the interaction sound, before returning to a fixed base level. Emotions no longer share one generic pitch glide. Each sound is designed according to the scene object and emotional direction. As the five-stage journey progresses, the sound gradually moves from thin and enclosed to more spatial and layered, developing alongside therapeutic text and visual transformation.