Hello, is it possible for the feature where the model able to copy the intonation of a speaking voice of one reference audio , and and use that intonation for other reference voice to generate the actual output?
kinda like what Index-tts2 has, cause right now, the model is just a bit too random and with out any emotion control that actually land.
thanks for the awesome work that you guys are providing btw!
Hello, is it possible for the feature where the model able to copy the intonation of a speaking voice of one reference audio , and and use that intonation for other reference voice to generate the actual output?
kinda like what Index-tts2 has, cause right now, the model is just a bit too random and with out any emotion control that actually land.
thanks for the awesome work that you guys are providing btw!