Sesame, the AI company behind the impressively realistic voice assistant Maya, has released the base AI model powering Maya, as it recently promised.
The model, which is 1 billion parameters in size (โparametersโ referring to individual components of the model), is under an Apache 2.0 license, meaning it can be used commercially with few restrictions. Called CSM-1B, the model generates โRVQ audio codesโ from text and audio inputs, according to Sesameโs description on the AI dev platform Hugging Face.
RVQ refers to โresidual vector quantization,โ a technique for encoding audio into discrete tokens called codes. RVQ is used in a number of recent AI audio technologies, including Googleโs SoundStream and Metaโs Encodec.
CSM-1B uses a model from Metaโs Llama family as its backbone paired with an audio โdecoderโ component. A fine-tuned variant of CSM powers Maya, Sesame says.
โThe model open-sourced here is a base generation model,โ Sesame writes in CSM-1Bโs Hugging Face and GitHub repositories. โIt is capable of producing a variety of voices, but it has not been fine-tuned on any specific voice [โฆ] The model has some capacity for non-English languages due to data contamination in the training data, but it likely wonโt do well.โ
Itโs unclear what data Sesame used to train CSM-1B. The company didnโt say.
The model has no real safeguards to speak of. Itโs an โhonor systemโ situation. Sesame is merely urging developers and users not to use the model to mimic a personโs voice without their consent, create misleading content like fake news, or engage in โharmfulโ or โmaliciousโ activities.
I tried the demo on Hugging Face, and cloning my voice took less than a minute. From there, it was easy to generate speech to my heartโs desire, including on controversial topics like the election and Russian propaganda:
Sesame, co-founded by Oculus co-creator Brendan Iribe, went viral in late February for its assistant tech, which comes close to clearing uncanny valley territory. Maya and Sesameโs other assistant, Miles, take breaths and speak with disfluencies, and can be interrupted while speaking, much like OpenAIโs Voice Mode.
Sesame has raised an undisclosed amount of capital from Andreessen Horowitz, Spark Capital, and Matrix Partners. In addition to building voice assistant tech, the company says itโs prototyping AI glasses โdesigned to be worn all dayโ thatโll be equipped with its custom models.


