FRAVE is a synthesis system, not a paper. The name is the method: fader networks applied to the RAVE architecture.
Deep generative audio models sound impressive and are almost impossible to play. The latent space that gives them their range is not a set of controls a musician can reach for; it is a coordinate system nobody asked for. FRAVE exists to make one playable.
This work makes that space playable. Salient musical features are explicitly removed from the latent representation using an adversarial confusion criterion, then reintroduced as separate conditioning information — so the feature you care about becomes an independent parameter rather than something entangled with everything else. The result behaves like a synthesiser knob: continuous, direct, and predictable in the direction it moves.
Because it stays small enough to embed, FRAVE is the engine behind the second generation of the Neurorack — the system moves off the bench and into a module a performer can patch. It was evaluated across instrumental, percussive and speech recordings, and supports both timbre transfer and attribute transfer.
The argument it makes is the same one running through the thesis and through the Absynth Preset Explorer: a model is only useful to a musician at the point where it becomes controllable.
Credits
- Role
- Model design, experiments, first author
- With
- Nils Demerlé Sarah Nabi David Genova Philippe Esling
- Support
- ACIDS group, IRCAM–STMS. Published at ICASSP 2023.