Portal & CoGlyssn2026Ɨ
IOSTESTFLIGHT Ā· ACTIVE BUILDCREATOR AND PRODUCT

Glyssn.

Camera-as-instrument iOS app — visual scenes composed into sound.

01 Ā· OVERVIEW

The brief.

PLATFORM
iOS
STATUS
TESTFLIGHT Ā· ACTIVE BUILD
ROLE
CREATOR AND PRODUCT

Glyssn turns the camera into a composition interface. Instead of generating music from a prompt, it maps readable visual inputs to musical parameters: hue sets key and scale, brightness controls tempo, motion drives rhythm density, and saturation shapes reverb. The product has two modes for two different user intents. Photo mode captures one framed moment and turns it into a short composition. Stream mode keeps composing as the user moves through a room. The design question was whether those mappings could feel predictable enough to be played, not just observed.

02 Ā· KEY ART
KEY ART
IN ONE LINE

Camera-as-instrument iOS app — visual scenes composed into sound.

03 Ā· MOBILE
Demo — stream mode
04 Ā· SCREENS
01 Ā· Stream mode — the camera is the instrument.

Stream mode — the camera is the instrument.

Stream mode — the camera is the instrument.

05 Ā· CHALLENGE

The challenge.

The challenge was not simply producing sound from video input. It was making the mapping feel intentional, learnable, and performance-ready. Users needed to understand cause and effect: what changes when they pan, hold still, move toward light, or point the camera at a saturated surface. The app also had to hold up outside ideal demo conditions. Low light, shaky hands, window glare, and low-motion scenes all affected the signal. The UX had to make the translation visible enough for users to trust it and the export flow polished enough for captured output to feel usable after the demo.

  • Real-time visual-to-audio mapping with perceptible cause and effect
  • Two distinct modes — single-moment capture vs continuous stream composition
  • Session persistence and one-tap stem export (WAV / MIDI)
  • Hold up in real environments, not just as a proof of concept
RESEARCH

Research.

I tested visual-to-audio mappings against a simple product standard: can a user predict what will change before they hear it? Hue to key, brightness to tempo, motion to rhythm, and saturation to reverb were selected because each input was understandable from the camera view. Early testing separated two use cases. Stream mode supports exploration and performance: the composition changes as the room changes. Photo mode supports capture and review: one frame becomes a saved piece. Keeping both required two completion models without making either mode feel secondary.

  • Color → key/scale, brightness → tempo, motion → rhythm, saturation → reverb
  • Stream mode rewards movement; photo mode rewards stillness and framing
  • Users need saved sessions — compositions should not vanish on close
  • Export must feel like a deliverable (stems), not a screen recording
Glyssn research
PROCESS

Process.

I defined the mapping logic, product architecture, and completion flow: aim, compose, review, export. The interface stayed intentionally minimal so the camera feed, mesh overlay, and audio response carried the interaction. AI tooling helped accelerate implementation, but the product decisions centered on legibility, latency, and export quality.

  • Live mesh overlay showing which visual channels drive which musical parameters
  • Capture Ā· play Ā· export loop across three core screens
  • Saved composition review beside an active stream session
  • Four-stem export — melody, pads, bass, rhythm as WAV or MIDI
DECISION 01

Stream Mode — Continuous Visual-to-Audio Mapping.

Stream mode is the live-performance model. As the user moves, hue shifts retune key, brightness changes tempo, motion changes rhythm density, and saturation controls reverb. The mesh overlay exposes the input layer so users can see what the camera is interpreting and learn how to play a space through movement.

  • Live parameter mesh tied to camera input
  • Continuous composition while moving
  • Readable cause-and-effect for learning and performance
DECISION 02

Photo Mode — Capture a Moment into a Finished Piece.

Photo mode changes the interaction from continuous performance to intentional capture. The user holds still, frames a scene, and Glyssn writes a short composition from that single image. It is the path for saving a specific visual moment as an audio artifact.

  • Single-frame capture → composed piece
  • Optimized for stillness and intentional framing
  • Same mapping engine, different completion trigger
Photo Mode — Capture a Moment into a Finished Piece
DECISION 03

Sessions & Stem Export.

Saved sessions turn the app from a live demo into a usable creative tool. Compositions persist after the camera moment ends, and export produces four stems: melody, pads, bass, and rhythm. WAV and MIDI output make the result portable for downstream production.

  • Saved sessions for stream and capture work
  • Review UI beside live stream for A/B listening
  • One-tap four-stem export
ITERATION

Iteration.

TestFlight moved the product beyond controlled prototypes. Real-room demos exposed weak spots: motion jitter in low light, tempo swings near bright windows, and export friction when users wanted something more useful than a screen recording. Each iteration tightened the link between visual input and audio output.

  • Live room demos surfaced latency and lighting edge cases
  • Stem export replaced screen-recording as the share path
  • Instagram demos became the primary way to show stream vs room behavior

Shipped TestFlight builds with stream and capture both stable enough to demo on camera, plus stem export for finished output.

Glyssn iteration
06 Ā· OUTCOME

The outcome.

Current State

Glyssn is live on TestFlight with stream mode, photo capture, saved sessions, stem export, and field demos recorded in real rooms.

What's Next

The next version expands the camera model from a frame to a room-scale instrument. A Bluetooth-connected Moog simulator companion is also in progress so camera-generated compositions can be shaped through hardware controls, not only on screen.

REFLECTION

Reflection.

What I Learned

Synesthesia as an interface only works when the mapping is legible. If users cannot predict what will change when they move the camera, the output reads as noise. The mesh overlay and coarse parameter ranges mattered more than adding musical complexity.

What I'm Most Proud Of

Getting the interaction to hold up live, in a hand, in a real room, while someone watches. TestFlight and the Instagram demos prove the loop works outside the studio.

Biggest Challenge

Balancing real-time responsiveness with output worth exporting. Stream mode needs immediacy; photo mode and stem export need polish. One mapping engine had to serve both without feeling like two separate apps.