@hn_036ac5
26 days ago
Perhaps not exactly "instantly generate a song based on a random blog post or group chat history", but more like "instantly generate a song based on an input prompt sentence" is suno.ai -- you should check it out!

Consumer AI music generator that turns a short text prompt into a finished song with vocals, lyrics, and instrumentation across a wide range of genres.
@hn_036ac5
26 days ago
Perhaps not exactly "instantly generate a song based on a random blog post or group chat history", but more like "instantly generate a song based on an input prompt sentence" is suno.ai -- you should check it out!
@hn_87f940
26 days ago
I just tried Suno and the results to me are terrible. It seems designed for making pop music no one will listen to. I have spent many hours with MusicLM making wild experimental music no one will listen to. MusicLM has no problem making really weird sound combinations. I just gave SunoAI some of my MusicLM prompts I have saved and the results are garbage. The problem with the AI test kitchen model though is the results sound like they are in mono. The ultimate for me will be when we can make rap/hiphop no one will listen to.
@hn_b197f9
27 days ago
> LLMs are so, so far from being able to the thinking that goes in a real-time musical improvisation context it's laughable. have you actually tried any of the commercial AI music generation tools from the last year, eg suno? not an LLM but rather (probably) diffusion, it made my jaw drop the first time i played with it. but it turns out you can also use diffusion for language models https://www.inceptionlabs.ai/
@hn_bf123e
28 days ago
wth that Suno AI site is amazing. How did i miss that? I imagine that Cafe's etc would love this.. they could tweak the music to suit their clientele and not have to pay any royalties.
@hn_4bab79
28 days ago
I don't have any special insight into how it works, but I suspect it is largely synthesizing audio from scratch. The more I've thought about it, the task of generating music feels very similar to the task of text-to-speech with realistic intonation. So feels like the same techniques would be applicable. Suno do have an open source repo here that presumably uses similar tech: https://github.com/suno-ai/bark > Bark was developed for research purposes. It is not a conventional text-to-speech model but instead a fully generative text-to-audio model, which can deviate in unexpected ways from provided prompts. Suno does not take responsibility for any output generated. Use at your own risk, and please act responsibly. I've generated probably >200 songs now with Suno, of which perhaps 10 have been any good, and I can't detect any pattern in terms of the outputs. Here's another one which is pretty good. I accidentally copied and pasted the prompt and lyrics, and it's amazing to me how 'musically' it renders the prompt: https://app.suno.ai/song/d7bad82b-3018-4936-a06d-8477b400aae... Here are a couple more which are pretty good (i use it primarily for making fun songs for my kids): https://app.suno.ai/song/a308ca8a-9971-47a3-8bb3-a95126ff1a8... https://app.suno.ai/song/3b78a631-b52a-4608-a885-94f2edc190b... And this one's kindof interesting in that it can render 'gregorian chant' (i mean it's not very good): https://app.suno.ai/song/0da7502b-73cf-4106-88e8-26f4f465a5f... But this is one reason it feels like these models are very similar to text-to-speech but with a different training set