Inflect-Micro-v2: complete voice in 9.36M parameters

(huggingface.co)

121 points | by nateb2022 8 hours ago

6 comments

  • yjftsjthsd-h 5 hours ago
    Couple highlights:

    > Complete local text-to-waveform speech synthesis under 10M parameters.

    In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.

    > English only, with one fixed male voice. This is not zero-shot voice cloning.

    (And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)

  • modinfo 3 hours ago
    This is amazing, the quality blow my mind for such small model! I just replaced my old onnx model with yours!

    here my implementation with speech dispatcher and server: https://github.com/skorotkiewicz/inflect-speechd

    thanks for shearing!

  • NetOpWibby 1 hour ago
    The inflections are weird but this doesn't sound like a robot. Not bad!
  • tmaly 6 hours ago
    This is impressive. I wish there were a voice clone option.
    • fastball 4 hours ago
      With so few parameters, I imagine a voice fine-tune might be readily tractable.
  • itake 2 hours ago
    Amazing quality for small size, but definitely not that enjoyable to listen to.

    IMHO, its at about the same quality level of historic TTS tools.

    • stavros 1 hour ago
      I'm not sure which historic tools you mean, but to me this sounds much better than anything older than ten years ago.
      • itake 35 minutes ago
        I compared the macos Samantha just now and I guess the inflect-micro is marginally better...
      • leobg 59 minutes ago
        Ivona „Joey“, „Amy“
  • jsomedon 5 hours ago
    amazing quality for such small size!