The agent writes notes, instruments and effects as text, renders, and reads the audio back as text: levels per bar, drum lanes, piano rolls, chords, song structure. Every song on this page was composed, sound-designed and mixed by an AI agent through ismail's tools, with a human listening and giving notes.
# beat, pitch, length, velocity notes_write track=bass bar=17 repeat=8 notes="0 A1 0.5 110; 0.75 A1 0.25 90" # a drum lane as 16th steps pattern_write track=kick bar=17 lanes='{"C1": "X...x.....X.x..."}'
render out=final mp3=also
rendered 78.8s -> renders/final.mp3
master: -11.8 LUFS, peak -0.3 dBFS
kick peak -2.4 rms -20.0 dB
bass peak -2.6 rms -16.3 dB
pad peak -9.6 rms -25.4 dB
master fx 0: max gain reduction -5.9 dB
Under each song is the text the agent used to judge it: the arrangement map (one character per bar, 9 = loudest, rows per frequency band, the bass note and the sections it found) and a piano roll of one part. Press play and the playhead runs across the map; click the map to jump to a bar.
Luigi Manson: the Luigi's Mansion theme as dubstep, made with ismail. One bass voice plays the hook, and the note velocity picks each hit's articulation: a vowel sweep, a wub, a screech, a talk word ("Mario?").
About 70 tools behind one MCP server (or a CLI). The agent never needs ears; it needs readings it can trust.
Notes, drum patterns, synth patches, effect chains, buses, automation. Instruments range from subtractive synths and drum synths to samplers and voices written as Python (the piano on this page is one: measured partials, no samples).
Tempo and grid, arrangement maps, levels and bands per bar, drum lanes as step strings, piano rolls, chords, timbre, vowels of a voice. A spectrogram PNG exists, but nothing needs it.
Give it a recording: it finds the grid, separates stems, transcribes only notes that repeat (echoes and bleed drop out), fits instruments to extracted sounds, and scores every draft against the reference's own variation, with drill-down to a single bar.
An agent skill teaches the workflow: plan the piece before writing notes, check every render with the analysis tools, and trust baseline-scored comparisons over a hunch.
Suno and similar services are models that turn a prompt into audio. ismail is the opposite kind of thing: a set of tools your own agent uses to write the song as notes, sounds and code, render it, read what came out, and edit. Here is how that compares with a music generator, and with driving a classic DAW like Ableton or FL Studio.
| ismail + your agent | Suno (music generator) | Ableton, FL Studio | |
|---|---|---|---|
| What makes the music | Your agent's own reasoning: harmony, rhythm, sound design and the math behind each synth | A trained model that generates audio from a prompt | You, with an agent at most pressing buttons through a bridge |
| How you change it | Edit one note, one patch, one bar, one fader; re-render in seconds; everything else stays exactly as it was | Regenerate, extend or replace sections of generated audio; stems and a browser studio on paid plans | Full manual control, by hand |
| What the agent can perceive | Every render comes back as text: levels per bar, drum lanes, piano rolls, chords, structure, scored comparisons | Nothing to perceive: you listen and write a new prompt | Usually nothing: the agent turns knobs without hearing the result |
| Sounds | Anything the agent can describe as code: synths, drum synths, samplers, voices written in Python, speech; any sound becomes an instrument | The palette the model learned | The plugins you own |
| The song is | A project file and a build script: git history, diffs, branches and code review for a track | Audio files | Binary project files |
| Runs | On your machine, open source (MIT), no content filter, no music subscription (you bring the agent: Claude Code, Cursor or any MCP client) | A hosted service with plans; the free plan is non-commercial | Paid licenses, closed source |
| Gets better when | The model driving it gets better: the same prompt, a stronger model, a better song | The vendor ships a new model | You get better |
| Extending it | Any agent or person can add an instrument, an effect or a way to listen, in a few lines of Python | Limited to what the service offers (including custom models trained on your own audio) | Plugin SDKs around a closed core |
They were designed before AI agents, for a person with ears and a mouse. Bridges and scripting let an agent press their buttons, but the instruments, channels and arrangement aren't laid out as text an agent can reason about, and the agent still can't hear what it did.
ismail starts from the other end: everything an agent needs to write and to perceive is compact text, and renders are fast. Its first working version was written by an AI agent in about a day, which is the point: if something is missing, your agent can add it.
Local does not mean copyright-free: a remix of someone else's melody is still their melody. What changes is that nothing here is generated by a model trained on other people's recordings; the notes and sounds are whatever your agent writes. Suno plan details as of 2026 (Suno).
Python 3.10+. Works with Claude Code, Cursor or any MCP client, or straight from the shell.
pip install git+https://github.com/newsbubbles/ismail claude mcp add -s user ismail -- python -m ismail.mcp_server
"make a 16 bar lo-fi loop in D minor in songs/demo and render an mp3"
Skill install, Cursor setup and the full tool list: README.