Summary
Separate query tools for the video tasks is the detail that says someone ran this in anger.
Video generation takes minutes, and a server that pretends otherwise times out halfway. Having explicit status queries for the image-to-video and text-to-video jobs means a long run can be started and collected rather than held open and lost. The suite is broad enough that the standing context cost is real — worth checking whether you need the whole thing or just the speech half.
What it is
SynClub's own server, bundling a wide range of generation backends — speech, voice, video, image — under a single connection.
What you get
- Speech — text to audio, voice cloning from a sample, and a dedicated Japanese path
- Video — text to video and image to video, with task-status queries for the long runs
- Image — generation, editing, recognition, background removal and high-definition restoration
- AI search
Requirements
Run with uvx. MIT licensed.
Setup effort
One command plus a key — uvx synclub-mcp, then supply credentials
