Multimodal
49 items across the graph — tagged with Multimodal.
From the graph · 49
Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience
🔍大模型应用开发实战一:RAG 技术全栈指南,在线阅读地址:https://datawhalechina.github.io/all-in-rag/
Open-source framework for building AI-powered apps in JavaScript, Go, and Python, built and used in production by Google
Solve Visual Understanding with Reinforced VLMs
A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
A Next-Generation Training Engine Built for Ultra-Large MoE Models
SimpleMem: Efficient Lifelong Memory for LLM Agents — Text & Multimodal
Open-source multimodal retrieval engine (Morphik Core). By Morphik — AI back office for skilled nursing & senior living (morphik.ai).
Hugging Face model with 3443 likes. Tags: gguf, uncensored, qwen3.6, moe, vision, multimodal, image-text-to-text, en, zh, multilingual
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
🤖 Type-safe, provider-agnostic TypeScript AI SDK for streaming chat, tool calling, agents, and multimodal apps across OpenAI, Anthropic, Gemini, React, Vue, Sv…
🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.
Implementation of "BitNet: Scaling 1-bit Transformers for Large Language Models" in pytorch
Let Claude (or any LLM) actually watch a video — scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.
Automated modeling and machine learning framework FEDOT
Your Open Autonomous Android Agent — A production-ready, self-planning AI assistant powered by local/remote LLMs and accessibility-driven screen automation.
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
Knowledge-Aware machine LEarning (KALE): accessible machine learning from multiple sources for interdisciplinary research, part of the 🔥PyTorch ecosystem. ⭐ St…
A curated archive of breakthroughs in Agents, Architecture, Training, RAG, and On-Device AI.
A plug-and-play compiler that delivers free-lunch optimizations for both inference and training.
Multimodal RAG to search and interact locally with technical documents of any kind
Multi-Agent System Powered by LLMs for End-to-end Multimodal ML Automation
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Self-hostable multimodal chat with local LLMs (Ollama/OpenAI): PDF RAG, image chat, and Whisper voice, Streamlit + Docker.
Enterprise-ready Spring AI platform for RAG, tool calling, async ingestion, JWT/RBAC security, and observability.
Corpus of resources for multimodal machine learning with physiological signals (mmps).
autoupdate paper list
100+ curated Seedance 2.5 prompts with real video previews, plus an installable Agent Skill that optimizes prompts, creates storyboards, and generates videos vi…
Efficient LLM inference on Slurm clusters.
VideoDB Python SDK
Open-source, self-hosted multi-agent AI infrastructure for teams and apps with support for cloud and local models, tool integrations, workflows, and automation…
An all in one AI solution compatible with any known AI service on the planet
The fastest way to put Volcengine Ark in your terminal and your AI agent — go from prompt to generated media, multimodal answer, or deployed endpoint in a singl…
Latest Advances on (RL based) Multimodal Reasoning and Generation in Multimodal LLMs
Production-grade multimodal RAG for financial document intelligence. Chart understanding · hybrid retrieval · numeric guardrails · multi-tenancy · full observab…
EmbodiedAgents is a fully-loaded ROS2 based framework for creating interactive physical agents that can understand, remember, and act upon contextual informatio…
All-in-One Multimodal Parsing Engine + Ontology-Powered AI-Ready Knowledge Engine Parse every modality. Compile knowledge with ontology. Reason before retrieval…
Unified interface for interacting with various LLMs hundreds of models, caching, fallback mechanisms, and enhanced reliability.
Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside embeddings, speech, and im…
Self-hosted personal AI agent and employee for workflow automation in your DMs. It writes code, runs tools, schedules jobs, saves workflows, and remembers conte…
🔥 Alternative to Ollama — multi-model serving with sub-ms model switching · CPU-only 20B inference for Edge AI · llama.cpp + stablediffusion.cpp
Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside embeddings, speech, and im…
Build use cases with VideoDB
A Go toolkit for building AI agents and applications across multiple providers. Unified LLM client, agent framework with handoffs, tool calling, streaming, stru…
Self-hosted AI knowledge base with hybrid semantic search (pgvector + FTS + RRF), MCP server, multi-provider LLM inference (Ollama, OpenAI, OpenRouter, llama.cp…
A minimal, hackable Vision-Language Model built on Karpathy’s nanochat — add image understanding and multimodal chat for under $200 in compute.
An open vision-language model for captioning and VQA.
Joint training recipes that align images and text in one embedding space.
A starter kit for training vision-language models.
