INTENTIO: Reclaiming the Digital Mind with PHP, Small Language Models, and Absolute Data Sovereignty
"Not a chatbot. Not a cloud service. A bounded space where AI pays attention."
1. The Story Behind INTENTIO: The Quiet Rebellion
The Illusion of the Cloud Oracle
The current artificial intelligence landscape has surrendered to a single narrative: that intelligence requires planetary scale, that useful cognition belongs exclusively in hyperscaler data centers, and that the only way to interact with an AI model is to stream your thoughts, proprietary data, and creative work across third-party networks to be logged, indexed, and monetized.
Modern AI products are designed as oracles. They are trained on the unfiltered entropy of the open internet. They attempt to know everything, yet they understand almost nothing about your specific domain, your private notes, your company’s trade secrets, or your intellectual boundary. When you query an oracle, you receive a generalized, smoothed-out consensus. Worse, you pay for that answer with your privacy, accepting the constant risk of hallucinations and the rent-seeking economics of cloud API tokens.
INTENTIO is a direct rejection of that consensus.
Created by Melasistema, INTENTIO began with a clear conviction: Software must run on hardware you own, over data that never leaves your machine, within systems you can read and inspect all the way down.
Private AI is not a toggle in a settings menu. It is not an enterprise compliance checkbox. Private AI is an architectural decision.
INTENTIO does not attempt to know everything. Instead, it creates a disciplined, bounded cognitive space where a Small Language Model (SLM) focuses entirely on your curated knowledge. It acts not as an omniscient oracle, but as a rigorous interpreter of meaning, strictly grounded in the facts and structures you define.
The PHP Anomaly: Why Write a Cognitive Engine in PHP?
In a developer culture intoxicated by multi-gigabyte Python virtual environments, fragile dependency trees, and ephemeral JavaScript runtimes, choosing PHP to build a modern cognitive framework sounds like heresy.
It is, in fact, an intentional design choice.
PHP has spent three decades refining one thing better than almost any other ecosystem on Earth: pragmatic data processing, filesystem manipulation, text parsing, and inspectable architecture.
- Modern PHP (8.2+) is fast, typed, rock-solid, and easily deployable.
- It does not suffer from virtual environment rot, broken system wheels, or opaque abstraction wrappers.
- Every byte of memory, every string transformation, every API interaction, and every filesystem operation can be read, debugged, and understood without magic.
Most local AI workflows do not need Python's deep learning training pipelines; they need orchestration. They need an intelligent, lean command-line engine that reads Markdown files, chunks documents, communicates over clean local protocols, coordinates vector representations, and formats prompts for inference.
PHP handles this orchestration with surgical precision and near-zero runtime overhead. In INTENTIO, PHP acts as the conductor, orchestrating high-performance local inference engines while keeping the application layer clean, inspectable, and human-readable.
The Evolution: Dropping Ollama in INTENTIO 0.3.0
For over a year, INTENTIO relied on Ollama as its primary inference backend. Ollama served the project well during its initial stages: it ran language models, generated vector embeddings, and, starting in version 0.2.1, briefly powered image generation workflows.
However, as INTENTIO pushed deeper into deterministic workflows and multimodal consistency, friction emerged:
- The Image Breakdown: Upstream changes in Ollama broke compatibility with the specific image models INTENTIO depended on (such as the Cartoon Universe character consistency package). The image pipeline was effectively stranded.
- Architectural Redundancy: Maintaining separate adapters for text, embeddings, and images within Ollama created unnecessary abstraction bloat.
- Closer to the Silicon: INTENTIO was developed on Apple Silicon. The macOS unified memory architecture and Neural Engine are among the most power-efficient local AI platforms available. Rather than running through general-purpose abstractions, INTENTIO needed a direct line to Apple’s MLX framework.
With version 0.3.0, INTENTIO dropped Ollama entirely.
Instead of a single heavy backend trying to do everything, INTENTIO split its architecture across two high-performance, MLX-native tools:
- oMLX (Text & Embeddings): A lightweight local server built natively on Apple MLX. It serves text generation and vector embeddings with true token-by-token streaming and batched embedding requests. The swap removed 309 lines of Ollama adapter boilerplate and added 309 lines of clean, dedicated oMLX client logic.
- mflux (Image Synthesis): A command-line tool built on Apple MLX running FLUX.2 klein 4B. Unlike background servers that sit idle consuming valuable VRAM,
mflux runs on demand, renders the image with live terminal progress, saves the output directly into the workspace, and exits cleanly. Crucially, mflux supports rendering from reference images, unlocking reproducible visual styles, consistent characters, and brand logo placement on physical mockups.
The result is a cognitive framework that is faster, leaner, and deeply integrated into local hardware.
┌────────────────────────────────────────────────────────┐
│ INTENTIO Core (PHP CLI) │
│ - Markdown Parser - Chunking & Ingestion Engine │
│ - Spaces & Packages - Prompt Templates & Intent │
└──────────────┬──────────────────────────┬──────────────┘
│ │
Text & Vector Embeddings On-Demand Images
▼ ▼
┌────────────────────────┐ ┌────────────────────────┐
│ oMLX │ │ mflux │
│ (Apple MLX Server) │ │ (Apple MLX CLI Tool) │
│ - SLM Text Generation │ │ - FLUX.2 klein 4B │
│ - Batch Embeddings │ │ - Image-to-Image / │
│ - Token Streaming │ │ Reference Consistency│
└────────────────────────┘ └────────────────────────┘
│ │
┌──────────────┴──────────────────────────┴──────────────┐
│ Apple Silicon Unified Memory (MLX) │
└────────────────────────────────────────────────────────┘
2. How INTENTIO Works: The Cognitive Architecture
INTENTIO operates on a foundational rule: Structure matters more than scale.
When you restrict a Small Language Model (typically between 3B and 8B parameters) to a clean, highly structured, well-embedded knowledge base, it frequently outperforms a 70B or 400B cloud giant that is struggling through a swamp of general internet noise.
Here is how INTENTIO guides local silicon through complex human knowledge:
[ Your Filesystem ]
spaces/domain/knowledge/
├── core-concepts.md
├── rules.md
└── context.md
│
▼ (./intentio ingest)
[ Intelligent Chunking & Header Preservation ]
│
▼ (Batched Vectorization)
[ oMLX Embedding Engine ]
│
▼
[ Local Vector Index (On-Device Storage) ]
│
├───────────────────────────────────────────────┐
│ User Query: ./intentio ask "..." │
▼ ▼
[ Semantic Similarity Search ] [ Markdown Prompt Template ]
│ │
└───────────────────────┬───────────────────────┘
│
▼
[ Grounded Context Assembly ]
│
▼
[ oMLX Local SLM ]
│
▼ (Real-Time Token Streaming)
[ Coherent, Unhallucinated Answer ]
Step 1: The Filesystem as a Semantic Map
In INTENTIO, you do not upload files to a database black box. Your filesystem is your knowledge map.
Knowledge is partitioned into Spaces located under spaces/. Inside each space, you organize your domain knowledge using plain Markdown files:
spaces/
my_brand/
knowledge/
01_brand_voice.md
02_value_propositions.md
03_audience_profiles.md
commands/
pitch.md
critique.md
manifest.json
Because the space is made of plain Markdown files, it remains 100% human-readable, trackable via Git, and editable in any text editor (Vim, Obsidian, VS Code, or Helix). You own the raw knowledge; INTENTIO merely interprets it.
Step 2: Ingestion & Fast Batched Embeddings
When you run:
./intentio ingest my_brand
INTENTIO parses your Markdown files, preserving structural hierarchies (headings, lists, code blocks). It divides content into semantically coherent chunks, retaining the document context inside the chunk metadata.
Under INTENTIO 0.3.0, chunks are embedded in a single batched request via oMLX. The entire space is vectorized in seconds on your Mac's Neural Engine, generating a fast, local index that resides safely inside your space folder.
Step 3: Grounded Retrieval Without Hallucination
When you ask INTENTIO a question:
./intentio ask my_brand "What is our position on third-party data tracking?"
INTENTIO does not let the model generate freely from its training data. Instead:
- It converts your query into a vector representation via
oMLX.
- It performs a cosine similarity search across your local index.
- It retrieves the exact sections of your Markdown documents that answer the query.
- It injects those sections into a strictly bounded prompt.
The Small Language Model is instructed to act purely as an interpreter of the provided context. If the answer is not in your documents, the model tells you. It does not fabricate sources, invent facts, or fill gaps with random web trivia.
Step 4: Declarative Command Templates
INTENTIO lets you define specialized tasks as Commands. A command is a Markdown template that lives in your space's commands/ directory.
For example, a critique.md command might instruct the model to adopt the persona of a skeptical investor, evaluate an incoming argument against your company's core tenets, and output a structured table of weaknesses.
You execute it effortlessly:
./intentio run my_brand:critique "We are thinking of adding a cloud telemetry dashboard"
The command template sets the tone, posture, and output format; the local space provides the truth; the SLM provides the reasoning.
Step 5: Native Token Streaming
In earlier versions, users waited in silence while the model completed its entire inference cycle before the terminal updated.
With INTENTIO 0.3.0's oMLX integration, tokens stream directly to your terminal the millisecond they are generated. You watch the model's reasoning unfold in real time at the native decode speed of your Apple Silicon hardware.
Step 6: On-Demand, Reference-Grounded Visuals with mflux
Generating images with local AI has historically meant either wrestling with complex Python ComfyUI webs or running heavy daemon servers that consume 12 GB of VRAM even when idle.
INTENTIO handles images via mflux with the ultra-efficient FLUX.2 klein 4B model:
- Zero Idle Footprint: INTENTIO spawns the CLI process only when you invoke an image command. Once the render completes, the memory is completely freed back to macOS.
- Live Progress Feedback: Terminal progress bars show step-by-step rendering progress.
- Visual Continuity via Reference Images:
mflux allows image generation conditioned on existing reference files. If you generate a character in step one, step two can reuse that character in a new scene. If you have a corporate vector logo, INTENTIO can stamp that logo onto a 3D product mockup without distorting the emblem.
3. Key Features
| Feature | What It Delivers | Why It Matters |
| Complete Data Sovereignty | 100% offline, on-device execution. Zero telemetry, zero cloud APIs, zero external requests. | Your intellectual property, code, private thoughts, and legal records never cross a network wire. |
| Inspectable PHP Core | Clean, modern PHP 8.2+ codebase with clear classes, zero framework bloat, and minimal dependencies. | You can open any file in the framework, understand how it works in 5 minutes, and adapt it to your needs. |
| Apple Silicon MLX Acceleration | Native integration with oMLX (text/embeddings) and mflux (FLUX.2 klein 4B images). | Maximizes M-series unified memory bandwidth and Neural Engine efficiency without Python runtime headaches. |
| Live Token Streaming | Immediate real-time output in the CLI. | Eliminates prompt latency anxiety; read the model's reasoning as it types. |
| Structured Markdown Knowledge | Knowledge bases and prompt commands are plain .md files in your filesystem. | No proprietary database lock-in. Compatible with Git, Obsidian, and standard UNIX text utilities. |
| Zero Idle Memory Drain | Images render via CLI execution rather than persistent background daemons. | Your laptop stays cool, your battery lasts, and RAM is freed the instant rendering finishes. |
| Reproducible Image Conditioning | Reference-image rendering support via mflux. | Visual consistency across renders—crucial for brand design, storytelling, and UI pitch mockups. |
| Modular Package System | Pre-packaged domain spaces installed with a single command (./intentio init <package>). | Start immediately with curated knowledge systems and tested commands out of the box. |
4. The Ecosystem in Action: Ready-Made Packages
INTENTIO is not an empty shell. It ships with ready-made Packages—self-contained spaces featuring curated knowledge bases, purpose-built Markdown commands, and tested prompt structures.
You can initialize any package into your local environment with one command:
./intentio init <package_name>
Here are three packages showcasing what bounded local cognition can do:
1. Hook Analyzer (hook_analyzer)
Writing effective copy usually involves guesswork or relying on generic ChatGPT platitudes. hook_analyzer turns INTENTIO into an objective copy auditor.
- The Knowledge Base: Encapsulates established psychological models, cognitive biases (curiosity gap, loss aversion, pattern interrupts), and rigorous platform-specific scoring rules.
- What It Does: Evaluates marketing hooks, scores them against cognitive retention metrics, contrasts two hooks side-by-side, and produces actionable platform reports (LinkedIn, X, newsletters, landing pages).
- The Sovereign Advantage: Test confidential marketing strategies, unpublished book titles, and commercial campaigns without feeding them into corporate training loops.
2. Cartoon Universe (cartoon_universe)
Visual storytelling requires persistent character design, recognizable silhouettes, and cohesive styling—something one-shot prompt engines routinely fail at.
- The Knowledge Base: Maintains strict style rules, color palettes, world-building lore, and aesthetic constraints.
- What It Does: Works alongside
mflux to generate stylized, expressive characters and scenes that adhere to unified aesthetic guidelines across multiple frames.
- The Sovereign Advantage: Artists and creators retain total control over character model sheets and artistic assets locally on their machines.
3. Product Pitch Lab (product_pitch_lab)
Creating marketing collateral for new products requires combining abstract messaging with concrete visual consistency.
- The Knowledge Base: Contains pitch structures (problem, solution, unique mechanism, proof), positioning frameworks, and landing page anatomy rules.
- What It Does: Refines your value proposition while using
mflux's reference image engine to place your authentic logo and product packaging into photorealistic scenes.
- The Sovereign Advantage: A founder can iterate from raw product concept to complete investor pitch deck and visual mockups in complete secrecy on a laptop during a flight.
5. Technical Requirements & Configuration
INTENTIO 0.3.0 is built deliberately for developer ergonomics on modern hardware.
System Prerequisites
- Hardware: Apple Silicon Mac (M1, M2, M3, M4, or later) with Unified Memory (16 GB+ recommended for running SLMs alongside image models).
- Runtime: PHP 8.2 or higher with standard extensions (
curl, mbstring, json).
- Text & Embedding Engine:
oMLX running locally.
- Visual Synthesis Tool:
mflux installed in your system PATH.
Quick Verification
Once installed, check your local environment directly from the CLI:
./intentio status
INTENTIO will inspect your local oMLX server, verify that the configured text model and embedding model are loaded, check that mflux is available, and flag any missing dependencies before you begin.
[✓] PHP Runtime: 8.3.4
[✓] oMLX Server: Connected (http://127.0.0.1:8000)
├── Text Model: Qwen2.5-7B-Instruct-4bit (Loaded)
└── Embedding Model: bge-large-en-v1.5 (Loaded)
[✓] mflux Engine: Available (/usr/local/bin/mflux)
└── Default Model: FLUX.2-klein-4B
[✓] Spaces Detected: 4 (my_brand, hook_analyzer, cartoon_universe, product_pitch_lab)
6. How Others Can Get Involved
INTENTIO is open-source, built in public, and dedicated to the philosophy of sovereign computing. We believe the future of AI will not belong to a handful of centralized cloud monopolies, but to an open confederation of local, personal, and domain-specific tools that respect the human mind.
There are many ways to participate in this quiet rebellion:
1. Build and Share Packages
The highest-leverage way to contribute is by authoring domain-specific Packages:
- Are you a lawyer? Build a legal research space containing statutory frameworks and compliance audit commands.
- Are you a software architect? Build a refactoring auditor loaded with Clean Architecture rules and design pattern guidelines.
- Are you a novelist? Create a narrative world-building space with character registries and continuity checkers.
- Any directory with Markdown knowledge files, a
commands/ folder, and a clean manifest.json can be packaged and shared with the community.
2. Contribute to the Core Engine
The INTENTIO codebase is deliberately clean and accessible. We welcome contributions on:
- Advanced Chunking Strategies: Improving semantic boundary detection for codebases, tabular Markdown, and multi-lingual texts.
- Vector Index Optimizations: Exploring even faster local search and indexing mechanisms.
- Command Tooling: Expanding CLI developer ergonomics, reporting tools, and interactive inspection modes.
- Alternative Local Backends: Refining client adapters and keeping pace with Apple MLX innovations.
3. Test and Benchmark Small Language Models
We are in a golden era of open-weights SLMs (Qwen 2.5, Llama 3.2, Gemma 2, Phi-3.5). Help test which quantization profiles and parameter sizes yield the cleanest reasoning inside bounded spaces. Share your findings and benchmark prompt recipes.
4. Star, Discuss, and Advocate
If you believe that individuals should own their intelligence rather than rent it from cloud vendors:
Clone the repository, inspect the code, point it at your private notes, and reclaim your digital mind.