Home / ML 101 / Module 12 / Lesson 2

Emerging Directions

Multimodal models, foundation models, AI agents, neuromorphic computing, quantum machine learning, and edge ML — the frontiers shaping the next decade of artificial intelligence.

~18 min read M12 · L2 Advanced

The Frontier Landscape

Machine learning is no longer a static discipline — it is an evolving field where fundamental assumptions shift every few years. The models of 2024 bear little resemblance to those of 2018, and the models of 2030 will likely differ just as dramatically. Six directions currently define the most active research frontiers: multimodal models, foundation models, AI agents, neuromorphic computing, quantum machine learning, and edge ML. Understanding these directions is essential not merely for staying current, but for building intuitions about where the field is heading and why.

These areas are not independent — they converge. A foundation model can power an AI agent; an edge ML chip may one day run a quantized multimodal model; quantum computing may ultimately accelerate the optimization underlying all of these. The unifying theme is capability expansion: each direction either extends what models can perceive, what they can do, where they can run, or how efficiently they compute.

Maturity Spectrum

These directions span a wide maturity range. Multimodal models and foundation models are production-deployed today. AI agents are rapidly transitioning from research to deployment. Edge ML is growing rapidly in embedded systems. Neuromorphic computing and quantum ML remain largely pre-commercial research frontiers, though with genuine long-term potential.

Multimodal Models

Traditional ML models operate on a single modality — images, text, or audio — with separate specialized architectures for each. Multimodal models process and generate across multiple modalities simultaneously, enabling richer understanding of the world. A model that can see an image, read its caption, and listen to accompanying audio understands more context than any single-modality system.

Key Architectures and Systems

GPT-4V / GPT-4o
Vision + Language
Extends GPT-4 with visual understanding. GPT-4o integrates audio, vision, and text end-to-end in a single model rather than via separate encoders.
Gemini 1.5
Native Multimodal
Designed from the ground up for multimodality. Processes text, images, audio, video, and code. Long context (1M tokens) enables processing of full video files or codebases.
DALL-E 3 / Stable Diffusion
Text → Image
Diffusion models conditioned on text embeddings. DALL-E 3 uses CLIP-based alignment; Stable Diffusion uses latent diffusion for efficiency.
Whisper
Audio → Text
OpenAI's speech recognition model trained on 680K hours of audio. Robust to accents, noise, and multiple languages; serves as the audio backbone in many multimodal pipelines.
CLIP
Vision-Language Alignment
Contrastive learning aligns image and text embeddings into a shared space. Zero-shot image classification, cross-modal retrieval, and foundation for most vision-language models.
Sora
Text → Video
Video generation via diffusion transformer on spatiotemporal patches. Demonstrates that language model scaling approaches generalize to video synthesis with temporal coherence.

The key technical challenge in multimodal models is modality alignment — ensuring that representations from different input types (image features, token embeddings, audio spectrograms) are mapped into a shared semantic space where relationships are meaningful. Contrastive pre-training (as in CLIP), cross-attention between modality-specific encoders, and unified tokenization (treating image patches as tokens) are the main approaches.

Emergent Multimodal Capabilities

Multimodal models exhibit capabilities that no single-modality model can: reading text embedded in images (OCR), understanding diagrams and charts, describing images in detail, answering questions about video frames, and generating images from complex compositional descriptions. These capabilities are not explicitly programmed — they emerge from scale and multimodal pre-training.

Foundation Models and Their Impact

The term foundation model (Bommasani et al., 2021) refers to large models trained on broad data at scale that can be adapted to many downstream tasks. GPT-4, Gemini, Claude, LLaMA, and BERT are all foundation models. What makes them foundational is not their size per se, but their generality: a single pre-trained model can be adapted, via fine-tuning or prompting, to tasks as varied as medical diagnosis, legal document analysis, and code generation.

Emergent Abilities

One of the most striking properties of foundation models is emergence: qualitative capabilities that appear abruptly at certain scales and are absent in smaller models. Few-shot in-context learning (the ability to solve tasks from a handful of examples in the prompt with no gradient updates), chain-of-thought reasoning, and multi-step arithmetic are all emergent behaviors observed in large language models beyond ~100B parameters — behaviors not explicitly optimized for during training.

PropertyDescriptionImplication
Scale100B–1T+ parameters trained on trillions of tokensEmergent capabilities unavailable in smaller models
TransferabilityPre-trained once, adapted many times via fine-tuning or promptingAmortizes training cost across thousands of downstream tasks
Few-shot LearningNew tasks solved from examples in context, no gradient updateRapid deployment to new domains without retraining
Homogenization RiskMany systems built on a single base modelErrors, biases, and failures propagate to all downstream systems
Data CentralizationTraining on internet-scale data requires enormous infrastructureAccess and development concentrated in large organizations

Neural Scaling Laws

Scaling laws (Kaplan et al., 2020; Hoffmann et al., 2022 — "Chinchilla") describe predictable power-law relationships between model size, dataset size, compute budget, and loss. The Chinchilla result showed that most large models were significantly undertrained: given a fixed compute budget, it is optimal to train a smaller model on more data rather than a larger model on fewer tokens. This led to models like Mistral-7B achieving GPT-3-level performance at a fraction of the parameters.

Neural Scaling Law
L(N) = \left(\frac{N_c}{N}\right)^{\alpha_N}
L(N) is the cross-entropy loss as a function of model parameters N. N_c is the critical scale, and α_N is the scaling exponent (empirically ~0.076 for language models). Loss decreases as a predictable power law with model size — enabling compute-optimal training decisions before running experiments.

AI Agents and Tool Use

A foundation model that only generates text is a powerful but passive tool. AI agents extend language models to take actions in the world: browsing the web, writing and executing code, calling APIs, querying databases, and chaining sequences of reasoning steps to accomplish multi-step goals. The agent paradigm shifts LLMs from passive question-answerers to active problem-solvers.

Core Agent Patterns

ReAct
Reason + Act
Interleaves reasoning traces ("Thought:") with tool calls ("Action:") and observations ("Observation:"). Enables transparent step-by-step problem solving with external tools.
Tool Calling
Function Invocation
Model outputs structured JSON to call predefined functions (search, calculator, code executor). Results are injected back into context. OpenAI function calling, Anthropic tool use.
RAG
Retrieval Augmented Generation
Retrieves relevant documents from a vector store at inference time and injects them into context. Grounds model responses in up-to-date factual information without retraining.
Plan-and-Execute
Multi-step Planning
Decomposes complex goals into subtasks, executes them sequentially or in parallel, and integrates results. Enables tackling long-horizon tasks beyond single-turn interactions.
Multi-Agent Systems
Collaboration
Multiple specialized agents (researcher, coder, critic) collaborate on tasks. AutoGen, CrewAI, and LangGraph implement structured multi-agent orchestration frameworks.
Memory Systems
Long-term Context
External key-value stores (Redis, Pinecone) persist information across sessions. Episodic memory retrieves relevant past interactions; semantic memory stores factual knowledge.

The central challenge in AI agents is reliability. A single model call that is 90% reliable becomes a system that fails 65% of the time after 10 sequential tool calls (0.9¹⁰ ≈ 0.35). Agent robustness requires error detection, retry logic, fallback strategies, and careful task decomposition to minimize irreversible actions. The field is actively developing evaluation benchmarks (WebArena, SWE-bench, GAIA) to measure real-world agent performance.

Neuromorphic Computing

Conventional deep learning runs on von Neumann architectures (CPUs and GPUs) that separate memory and computation, creating a memory bandwidth bottleneck for large matrix multiplications. Neuromorphic computing takes inspiration from biological neural systems to build hardware where computation and memory are co-located — neurons in neuromorphic chips both store and compute in place, eliminating the memory-compute bottleneck.

Spiking Neural Networks

Spiking neural networks (SNNs) are the primary computational model for neuromorphic hardware. Unlike artificial neurons that output continuous activations, spiking neurons fire discrete binary events (spikes) at specific times. Information is encoded in the timing and rate of spikes rather than real-valued activations. SNNs are inherently event-driven: a neuron consumes energy only when it fires, in contrast to dense matrix multiplications which consume energy even when inputs are zero.

Intel Loihi 2
Research Neuromorphic Chip
128 neuromorphic cores, 1M programmable neurons, on-chip learning. ~1000× more energy efficient than GPU for sparse SNN workloads. Used for robotics, adaptive sensing, and edge inference.
IBM NorthPole
Near-memory Inference
256 compute cores with on-chip SRAM. Eliminates off-chip DRAM access for inference. Achieves 25× better energy efficiency than GPU A100 on ResNet-50 benchmark.
BrainScaleS
Analog Neuromorphic
Mixed-signal (analog + digital) hardware from University of Heidelberg. Runs 10,000× faster than biological time via analog computation. Used for neuroscience simulation and SNN research.
SpiNNaker 2
Massively Parallel
Manchester's second-generation neuromorphic system. 152 ARM cores, energy-proportional computation. Designed for real-time processing of sensor data streams with extremely low power.

The primary challenge in neuromorphic computing is the training problem: backpropagation does not directly apply to SNNs because the spike function is non-differentiable. Surrogate gradient methods approximate the spike function with a smooth surrogate during backpropagation. Alternatively, SNNs can be created by converting pre-trained ANNs (ANN-to-SNN conversion). The resulting models typically sacrifice some accuracy relative to their ANN counterparts in exchange for dramatic power savings — making neuromorphic systems most attractive for always-on, battery-powered edge applications.

Quantum Machine Learning

Quantum machine learning (QML) explores how quantum computers can accelerate or qualitatively enhance machine learning algorithms. Quantum computers represent information as qubits, which unlike classical bits can exist in superpositions of 0 and 1 simultaneously. This superposition, combined with entanglement and interference, allows quantum algorithms to explore exponentially large solution spaces in ways classical computers cannot efficiently replicate.

Qubit State
|\psi\rangle = \alpha|0\rangle + \beta|1\rangle, \quad |\alpha|^2 + |\beta|^2 = 1
A qubit exists in superposition of |0⟩ and |1⟩ with complex amplitudes α and β. The normalization constraint |α|² + |β|² = 1 ensures total probability sums to one. Measurement collapses the qubit to |0⟩ with probability |α|² or |1⟩ with probability |β|². An n-qubit system represents 2ⁿ states simultaneously — enabling parallel exploration of exponentially large state spaces.

Variational Quantum Circuits

The most practical near-term approach to QML is variational quantum circuits (VQCs), also called quantum neural networks. A VQC applies a parameterized sequence of quantum gates to qubits, then measures the output. Parameters are optimized classically using a hybrid quantum-classical loop — the quantum circuit computes the forward pass, and gradients are estimated via the parameter-shift rule or finite differences, then classical optimizers update parameters. VQCs have been demonstrated for classification, regression, and generative tasks on small quantum hardware.

QML ApproachQuantum Advantage ClaimedCurrent Status
Quantum SVM (QSVM)Exponential speedup for kernel evaluation in high-dimensional feature spaceTheoretical; dequantization results reduce advantage for classical data
Quantum PCA (qPCA)Exponential speedup in eigenvalue decomposition for density matricesRequires quantum RAM (QRAM) which is far from practical
Variational QNNMay learn functions hard for classical networks on quantum-native dataDemonstrated on small datasets; barren plateau problem limits trainability
Quantum Boltzmann MachineSample complex distributions more efficiently using quantum annealingD-Wave annealers used; limited to low-connectivity graphs
The Dequantization Caveat

Many claimed quantum speedups for ML have been "dequantized" — classical algorithms (sampling, randomized linear algebra) were subsequently found to achieve similar asymptotic complexity. The case for genuine quantum advantage in ML remains strongest for problems where the input data is itself quantum (quantum chemistry, quantum simulation) rather than classical data encoded into quantum states.

Edge ML and TinyML

Deploying machine learning at the edge — on microcontrollers, smartphones, IoT sensors, and embedded systems — imposes severe constraints: milliwatts of power, kilobytes of RAM, megabytes of flash storage, and milliseconds of latency. TinyML is the discipline of making neural networks run within these constraints without cloud connectivity. This is essential for applications like always-on keyword spotting, on-device health monitoring, smart cameras, and autonomous vehicles where cloud round-trips introduce unacceptable latency or privacy risk.

Model Compression Techniques

Quantization
Reduce numerical precision
Convert 32-bit floats to 8-bit integers (INT8) or 4-bit integers (INT4). 4× memory reduction with minimal accuracy loss. Post-training quantization (PTQ) or quantization-aware training (QAT).
Pruning
Remove redundant weights
Zero out weights below a threshold (unstructured) or remove entire filters/heads (structured). Structured pruning yields actual speedups on hardware; unstructured requires sparse execution support.
Knowledge Distillation
Teacher-student compression
Train a small student model to mimic the soft output distributions of a large teacher. The student learns richer representations than training on hard labels alone. Used for DistilBERT, TinyBERT.
MCUNet
NAS for microcontrollers
MIT's neural architecture search framework co-designs network architecture and inference engine for MCUs with <1 MB SRAM. Achieves ImageNet accuracy on microcontrollers for the first time.
MobileNets / EfficientNets
Efficient architectures
Depthwise separable convolutions (MobileNet), compound scaling (EfficientNet). Designed from scratch for inference efficiency rather than compressed from large models.
TFLite / ONNX Runtime
Inference runtimes
Optimized inference frameworks for edge hardware. Operator fusion, memory planning, hardware-specific kernels (ARM NEON, DSP). Deploy to phones, embedded Linux, MCUs.

A key challenge in TinyML is that the most accurate models are the largest, while edge constraints demand small models. Neural architecture search (NAS) automates the design of architectures that hit a specific Pareto frontier of accuracy versus model size or latency — searching over billions of possible network configurations to find those that fit within given hardware budgets. Hardware-aware NAS (like MCUNet and Once-for-All) directly incorporates hardware latency tables into the search objective.

Convergence: The Unified Future

The most important insight about these emerging directions is how they converge. Foundation models provide the reasoning backbone for AI agents; multimodal encoders extend what agents can perceive; edge ML brings agents to resource-constrained environments; and neuromorphic hardware may eventually power edge agents with biological-level energy efficiency. Quantum computing, though more speculative, may ultimately accelerate the optimization, sampling, or inference steps underlying all of these systems.

Consider a near-future autonomous vehicle: a quantized multimodal foundation model (vision + LIDAR + map) runs on a neuromorphic-inspired inference chip at the edge, with an agentic planning layer that calls external navigation and traffic APIs, all updated via federated learning without transmitting raw sensor data. Every emerging direction described in this lesson plays a role in that system.

Outstanding Challenges


Key Takeaways

Multimodal models align vision, language, and audio into shared representations (CLIP contrastive training; diffusion for generation). Foundation models exhibit emergent abilities at scale, governed by power-law scaling laws — the Chinchilla result shows compute-optimal training uses smaller models on more data. AI agents extend LLMs to take actions via tool calling, RAG, and multi-step planning (ReAct). Neuromorphic chips use spiking neurons for event-driven, energy-proportional computation (Intel Loihi, IBM NorthPole). Quantum ML uses variational circuits in hybrid quantum-classical loops, with genuine advantage clearest for quantum-native data. Edge ML deploys compressed models (quantization, pruning, distillation, NAS) to microcontrollers and phones. All six directions converge toward capable, efficient, autonomous ML systems that operate across the full compute spectrum.

Previous M12-L1: Responsible ML Module Overview Next Lesson M12-L3: Career and Resources