Overseas access: www.kdjingpai.com

Bookmark Us

Multimodal real-time interactive products

 Submit Website

TEN: An open source tool for building real-time multimodal speech AI intelligences
TEN Framework is an open source software platform focused on helping developers build real-time, multimodal, low-latency speech AI intelligences. It supports multiple programming languages including C, C++, Go, Python, JavaScript, and TypeScript.Developers can use the TEN Framework to quickly create speech, visual, and text...
07-30 2.3 K0kudos
wukong-robot: a smart speaker project to create personalized Chinese voice conversations
wukong-robot is an open source Chinese voice conversation robot and smart speaker project, designed to help developers quickly build personalized smart speakers. It supports Chinese speech recognition, speech synthesis and multi-round dialog features , integrated with ChatGPT, Baidu, KDDI and other technologies. The project design is modular, plug-ins and features can be freely extended, suitable...
07-24 2.4 K0kudos
BAGEL
BAGEL is an open source multimodal base model developed by the ByteDance Seed team and hosted on GitHub.It integrates text comprehension, image generation, and editing capabilities to support cross-modal tasks. The model has 7B active parameters (14B parameters in total) and uses Mixture-of-Tra...
05-22 3.3 K0kudos
RealtimeVoiceChat
RealtimeVoiceChat is an open source project that focuses on real-time, natural conversations with artificial intelligence via voice. Users use the microphone to input speech, the system captures the audio through the browser, quickly converts it to text, generates a reply from a large language model (LLM), and then converts the text to speech output, the whole process is close to real-time. The project adopts ...
05-06 4.2 K0kudos
Stepsailor: Integrating AI Command Bars in Existing SaaS Offerings
Stepsailor 是一个专为开发者打造的工具，核心是一个 AI 命令栏。开发者可以用它让自己的软件产品听懂用户的话，比如用户说“添加新任务”，软件就自动执行。它通过简单的 SDK 集成到 SaaS 产品中，不需要开发者懂 AI 技术。S...
04-10 2.2 K0kudos
OpenAvatarChat: modularly designed digital human conversation tool
OpenAvatarChat is an open source project developed by the HumanAIGC-Engineering team and hosted on GitHub. It is a modular digital human conversation tool that allows users to run full functionality on a single PC. The project combines real-time video, speech recognition, and digital human technology...
04-05 4.1 K0kudos
VideoMind
VideoMind is an open source multimodal AI tool focused on inference, Q&A and summary generation for long videos. It was developed by Ye Liu of the Hong Kong Polytechnic University and a team from Show Lab at the National University of Singapore. The tool mimics the way humans understand video by breaking down the task into steps such as planning, positioning, verifying and answering, one by...
04-02 3.4 K0kudos
MoshiVis
MoshiVis is an open source project developed by Kyutai Labs and hosted on GitHub. It is based on the Moshi speech-to-text model (7B parameters), with about 206 million new adaptation parameters and the frozen PaliGemma2 visual coder (400M parameters), allowing the model...
03-28 3.2 K0kudos
Qwen2.5-Omni
Qwen2.5-Omni is an open source multimodal AI model developed by Alibaba Cloud Qwen team. It can process multiple inputs such as text, images, audio, and video, and generate text or natural speech responses in real-time. The model was released on March 26, 2025, and the code and model files are hosted on GitHu...
03-27 4.9 K0kudos
xiaozhi-esp32-server: Xiaozhi AI chatbot open source back-end services
xiaozhi-esp32-server is a tool to provide backend service for Xiaozhi AI chatbot (xiaozhi-esp32). It is written in Python and based on the WebSocket protocol to help users quickly build a server to control ESP32 devices. This project is suitable ...
03-18 9.7 K0kudos
Baichuan-Audio
Baichuan-Audio is an open source project developed by Baichuan Intelligence (baichuan-inc), hosted on GitHub, focusing on end-to-end voice interaction technology. The project provides a complete audio processing framework that can convert speech input into discrete audio tokens , and then through a large model to generate the corresponding text ...
02-28 2.9 K0kudos
PowerAgents: AI Intelligent Body Platform for Timing Web Tasks
PowerAgents is an AI intelligences platform focused on web automation tasks, which allows users to create and deploy AI intelligences capable of clicking, entering and extracting data. The platform supports setting tasks to run automatically on an hourly, daily or weekly basis, and users can watch the intelligences at work in real time. It not only provides autonomous building capabilities, but also has social...
02-28 2.5 K0kudos
Step-Audio
Step-Audio is an open source intelligent voice interaction framework designed to provide out-of-the-box speech understanding and generation capabilities for production environments. The framework supports multi-language dialog (e.g., Chinese, English, Japanese), emotional speech (e.g., happy, sad), regional dialects (e.g., Cantonese, Szechuan), and adjustable speech rate and rhythmic style (e.g., rap). step-...
02-19 3.1 K0kudos
Gemini Cursor: an AI desktop smart assistant built on Gemini that can see, hear and speak
Gemini Cursor is a desktop intelligent assistant based on Google's Gemini 2.0 Flash (experimental) model. It enables visual, auditory, and voice interactions via a multimodal API, providing a real-time, low-latency user experience. The project, created by @13point5, aims to pass...
02-12 2.9 K0kudos
DeepSeek-VL2
DeepSeek-VL2 is a series of advanced Mixture-of-Experts (MoE) visual language models that significantly improve the performance of its predecessor, DeepSeek-VL. The models excel in tasks such as visual quizzing, optical character recognition, document/table/diagram comprehension, and visual localization.De...
02-12 3.5 K0kudos
AI Web Operator: Browser Automation, an Open Source Implementation of OpenAI Operator
AI Web Operator is an open source AI browser operator tool designed to simplify the user experience in the browser by integrating multiple AI technologies and SDKs. Built on Browserbase and the Vercel AI SDK, the tool supports a variety of Large Language Models (LLM)...
01-31 3.1 K0kudos
SpeechGPT 2.0-preview: an end-to-end anthropomorphic speech dialog grand model for real-time interaction
SpeechGPT 2.0-preview is the first anthropomorphic real-time interaction system introduced by OpenMOSS, which is trained on millions of hours of speech data. SpeechGPT 2.0-previ...
01-30 2.8 K0kudos
OpenAI Realtime Agents
OpenAI Realtime Agents is an open source project that aims to show how OpenAI's real-time APIs can be utilized to build multi-intelligent body speech applications. It provides a high-level intelligent body model (borrowed from OpenAI Swarm) that allows developers to build complex multi-intelligent body speech systems in a short period of time. The project ...
01-19 3.6 K0kudos
Bailing
Bailing (Bailing) is an open-source voice conversation assistant designed to engage in natural conversations with users through speech. The project combines speech recognition (ASR), voice activity detection (VAD), large language modeling (LLM), and speech synthesis (TTS) technologies to implement a GPT-4o-like voice conversation bot. The end-to-end latency of BaiLing's ...
01-19 3.4 K0kudos