An AI-driven systems engineering orchestration pipeline that bridges deterministic physical modeling (MATLAB/Simulink) with probabilistic reasoning (Large Language Models).
This tool acts as a "Systems Engineering Copilot," allowing engineers to semantically query complex Simulink architectures, Stateflow finite state machines (FSMs), and hardware deployment configurations through an interactive conversational interface.
- Topological Graph Mapping: Extracts not just disconnected blocks, but the actual Directed Acyclic Graph (DAG) by evaluating
PortConnectivityedges to map signal flows across the entire system. - Hardware Support Package Awareness: Dynamically queries custom
MaskNamesandMaskValuesto extract specific configurations for embedded hardware blocks (e.g., Arduino Pin mappings, ROS2 nodes) rather than falling back to genericMATLABSystemtypes. - Global Configuration Context: Captures the deployment environment, including Solver settings, Hardware Board targets, and MATLAB Workspace lifecycle callbacks (
InitFcn,StopFcn).
- Uses MATLAB's
sfrootAPI to penetrate Stateflow charts. - Extracts internal variables, hierarchical states (with Entry/During/Exit actions), and strict transition logic (Conditions, Triggers, and Transition Actions).
- Converts the highly nested JSON Abstract Syntax Tree (AST) into clean, LLM-optimized Markdown using a deterministic Jinja2 templating engine.
- Employs structural Markdown chunking to ensure data tables and state machine logic are never fractured during vector embedding.
- Vector Store: In-memory 3072-dimensional ChromaDB powered by the 2026-generation
gemini-embedding-001model. - Reasoning Engine: LangChain orchestration utilizing
gemini-2.5-flash, governed by a strict prompt enforcing rigorous systems engineering terminology and preventing physical parameter hallucination.
- Backend: Non-blocking FastAPI/Uvicorn server utilizing the
lifespanpattern to hold the LLM and ChromaDB in memory, delegating LangChain I/O to background threadpools. - Frontend: A sleek, interactive Streamlit chat interface for real-time querying.
- Extraction Layer (
src/core/): Headless MATLAB Engine parsing Simulink/Stateflow topologies into strict Pydantic V2 schemas. - Generation Layer (
src/generator/): Jinja2 rendering engine for deterministic Markdown formatting. - Intelligence Layer (
src/agent/): LangChain RAG architecture integrating ChromaDB and Google GenAI APIs. - Service Layer (
src/api/&app.py): Asynchronous FastAPI backend coupled with a Streamlit chat frontend.
- Language: Python 3.11+, MATLAB R2024b
- AI/ML: LangChain, Google GenAI API, ChromaDB
- API/Web: FastAPI, Uvicorn, Streamlit
- Validation: Pydantic V2
1. Environment Setup Ensure MATLAB R2024b is installed and properly licensed. Clone the repository and install the dependencies:
conda create -n doc-agent python=3.11
conda activate doc-agent
pip install -r requirements.txt2. Configure API Keys Copy the environment template and add your Google API key:
cp .env.example .env3. Boot the Pipeline
Run the root orchestrator. This will prompt you to select a .slx file, extract the AST, and boot the FastAPI backend on localhost:8000.
python main.py4. Launch the Client Interface
streamlit run app.py