AI directory search

Search across educators, skills, and resources.

Use this when you know the topic you need: Claude Code, MCP, evals, RAG, agents, product, coding, prompting, foundations, or model internals.

22 matches for "multimodal"

Video matches

Watch first when you want a fast feel for the topic before opening courses, docs, or profiles.

Google's AI endgame is here... everything you missed at I/O 2026 video thumbnail

Google's AI endgame is here... everything you missed at I/O 2026

Fireship · 2026, gemini, frontier models, multimodal ai

Providers and platforms

Official Gemini material for learning Gemini 3.8 Flash and the current stable, preview, latest, and experimental model lineup, plus the now-GA Interactions API, Lyria music generation, prompt design, function calling, background execution, Deep Research, Computer Use, Hooks, Live API, File Search, coding-agent setup, multimodal tradeoffs, and AI Studio workflows.

Topics

Gemini 3.8 Flash, Gemini models, Multimodal AI, Long context, Model selection, AI Studio, Interactions API, Background execution, Deep Research, Computer Use, Hooks, Live API, Coding agents, Music generation, File Search, API examples

Mistral AI profile photo

Mistral AI

Mistral models docs · Beginner to advanced

Official material for comparing current Mistral families such as Devstral, Magistral, Voxtral, OCR, and newer Small and Medium models across coding, reasoning, and multimodal use cases.

Topics

Mistral models, Open models, Model selection, Agents, Coding, Reasoning, Multimodal AI

Qwen profile photo

Qwen

Qwen docs · Beginner to advanced

Official Qwen material for learning the Qwen3 family, multilingual and multimodal capabilities, local deployment paths, and function-calling behavior.

Topics

Qwen, Open models, Multilingual AI, Coding agents, Multimodal AI, Model selection

Resources

Gemini 3.8 Flash

Model docs · Google AI for Developers · Intermediate

You need Google's official September 2, 2026 GA reference for its newest Flash model, including the stable model ID, 1M-token input window, thinking levels, multimodal inputs, tools, and production inference options.

gemini, gemini 3.8 flash, coding, autonomous agents, long context

Lyria 3.5 music generation

Model docs · Google AI for Developers · Beginner to intermediate

You want Google's September 3, 2026 preview documentation and examples for generating short clips or full songs from text and image inputs with the Lyria 3.5 Clip and Pro models.

google, lyria 3.5, music generation, audio generation, multimodal

Multi-Vector Embedding Models with Sentence Transformers

Practical guide · Hugging Face · Intermediate to advanced

You want a runnable guide to ColBERT-style late-interaction retrieval, including MaxSim scoring, retrieve-and-rerank, indexing, visual document retrieval, evaluation, and the storage-quality tradeoff versus dense embeddings.

hugging face, sentence transformers, rag, retrieval, embeddings

Cohere Parse

Model docs · Cohere · Intermediate

You need Cohere's August 27, 2026 document-parsing model for turning PDFs, slides, forms, tables, and images into structured Markdown before embedding, reranking, or agent retrieval.

cohere, parse, document ai, rag, ocr

Gemini API models

Model docs · Google AI for Developers · Beginner to advanced

You need to compare current Gemini stable, preview, latest, and experimental model IDs, context windows, and modality support.

gemini, google, multimodal, long context, model selection

Gemini Live API overview

Guide · Google AI for Developers · Intermediate

You want the official Google path for low-latency voice and vision agents before designing a realtime Gemini workflow.

gemini, live api, voice agents, realtime, multimodal

Gemini prompt design strategies

Guide · Google AI for Developers · Intermediate

You want current Gemini-specific prompting advice for system instructions, reasoning tasks, multimodal inputs, and tool-connected workflows.

gemini, prompting, reasoning, multimodal, tool use

Llama 4 prompt formats

Model docs · Meta Llama · Beginner to advanced

You need the current official Llama 4 prompt formats and model-card details before choosing Maverick or Scout for hosted or local deployment.

llama, meta, llama 4, prompt formats, open models

Mistral models overview

Model docs · Mistral AI · Beginner to advanced

You need to compare current Mistral families such as Devstral, Magistral, Voxtral, OCR, and general-purpose models.

mistral, open models, commercial models, model selection, coding

DeepSeek-V4-Flash-Vision-Exp Release

Release notes · DeepSeek · Intermediate

You want the Friday, August 21, 2026 DeepSeek launch note for the new multimodal V4 Flash variant, including image support, Files API reuse, and agent-workflow implications.

deepseek, multimodal, vision, agents, responses api

DeepSeek Vision

Guide · DeepSeek · Intermediate

You need the operational guide for image inputs, token usage, supported formats, and multimodal prompting after the August 21, 2026 launch of `deepseek-v4-flash-vision-exp`.

deepseek, vision, multimodal, agents, ocr

Qwen Model Studio model list

Model catalog · Qwen · Intermediate

You need the current hosted Qwen and third-party model catalog with modality coverage and capability splits.

qwen, model selection, multimodal, reranking, embeddings

Gemini Interactions API

Guide · Google AI for Developers · Intermediate to advanced

You need Google's current recommended API surface for agentic workflows, server-side state, and complex multimodal conversations.

gemini, interactions api, agents, tool use, stateful workflows

Gemini File Search

Guide · Google AI for Developers · Intermediate

You want Google's newest first-party retrieval path for grounded Gemini answers with managed chunking, indexing, and multimodal embeddings.

gemini, file search, rag, interactions api, grounded answers

Mistral model selection guide

Guide · Mistral AI · Beginner to advanced

You want Mistral's official comparison of model families, pricing, context, and licensing before implementation.

mistral, model selection, open models, coding, multimodal

Mistral OCR guide

Guide · Mistral AI · Intermediate

You want Mistral's current document-AI path for OCR, PDF extraction, and downstream RAG or automation workflows instead of treating OCR as a separate stack.

mistral, ocr, document ai, multimodal, extraction

Mistral OCR 4.1

Model docs · Mistral AI · Intermediate

You want the current OCR 4.1 model page before building document parsing, block-level extraction, or PDF-to-RAG workflows on older Mistral OCR assumptions.

mistral, ocr, document ai, bounding boxes, confidence scores

Cohere Command A Vision

Model docs · Cohere · Intermediate

You want Cohere's current multimodal Command model for chart reading, document OCR, and image-grounded enterprise workflows instead of assuming Command is text-only.

cohere, command a vision, vision, documents, ocr