The software industry is undergoing a generational transition from conversational artificial intelligence to agentic, programmatic tool use. For the past three years, the dominant paradigm of generative AI has been conversational chat interfaces: a human inputs natural language, and a large language model (LLM) returns streaming markdown text. While conversational interfaces are impressive for drafting essays, generating boilerplate code, or summarizing text, they remain fundamentally passive. They cannot autonomously alter real-world data or execute multi-step operational workflows without human intervention.
To bridge this gap, modern AI engineering has embraced agentic tool planning. However, nearly all existing agent frameworks—such as LangChain, AutoGen, CrewAI, and Semantic Kernel—rely heavily on server-side Python runtimes, Celery queue workers, and remote Docker containers. These architectures introduce severe infrastructure overhead, latency bottlenecks, and profound data privacy vulnerabilities when manipulating user files.
When designing NexaBatch, we asked a radical engineering question: Can we build a fully deterministic, multi-tool AI planning engine that runs entirely inside a standard client web browser, coordinates over 80 independent media and developer tools, and requires zero backend server infrastructure?
This article provides an in-depth technical analysis of the NexaBatch architecture. We explore how we utilize Google's Gemini 3.6 Flash model for sub-second structured Directed Acyclic Graph (DAG) planning, how we dynamically load heavy WebAssembly dependencies just-in-time, how we manage browser memory across zero-copy binary streams, and how our client-side fallback system guarantees resilience against edge-case failures.
The Planner-Executor Architectural Pattern§
At the core of NexaBatch lies a decoupled Planner-Executor design pattern. Rather than allowing a non-deterministic AI model to directly execute arbitrary JavaScript code or eval untrusted scripts in the browser, the system enforces a strict structural boundary between planning and execution.
┌────────────────────────────────────────────────────────────────────────────┐
│ NEXABATCH DUAL-ENGINE ARCHITECTURE │
└────────────────────────────────────────────────────────────────────────────┘
[User Natural Language Prompt]
│
▼
┌──────────────────────────────────────────────────────────┐
│ CONTROL PLANE: Gemini 3.6 Flash Planner │
│ • Receives: Prompt + File Metadata (MIME, Size, Name) │
│ • Context: 80+ Compact JSON Tool Schemas │
│ • Configuration: Temperature 0.1, JSON Schema Mode │
│ • Output: Strictly Typed JSON Execution DAG │
└────────────────────────────┬─────────────────────────────┘
│
(Deterministic JSON DAG)
│
▼
┌──────────────────────────────────────────────────────────┐
│ DATA PLANE: In-Browser Client-Side Executor │
│ • Dependency Resolver: JIT Lazy CDN Script Ingestion │
│ • Type Validator: Checks input/output MIME compatibility │
│ • Memory Stream: Zero-copy Blob[] and ArrayBuffer pipe │
│ • Sandboxed Engines: WebAssembly, WebCodecs, Canvas, │
│ Web Crypto API (window.crypto.subtle) │
└────────────────────────────┬─────────────────────────────┘
│
▼
[Final Processed Files]
The Control Plane vs. The Data Plane§
- The Control Plane (The Planner): The planning engine is responsible exclusively for reasoning, semantic intent extraction, and topological dependency ordering. It receives user intent (e.g., "Take these RAW screenshots, crop out the taskbar, convert them to AVIF, and generate a hash manifest") along with shallow metadata about the incoming files. It produces a clean, validated JSON execution manifest detailing which tool adapters must execute, in what order, and with what specific configuration arguments.
- The Data Plane (The Executor): The execution engine is completely local, deterministic, and sandboxed. It takes the JSON manifest produced by the Planner, validates MIME-type compatibility across adjacent steps, resolves and loads necessary external script dependencies on demand, and executes each tool adapter sequentially in browser RAM.
By strictly separating the Control Plane from the Data Plane, NexaBatch achieves three critical advantages:
- Absolute Security: The LLM never touches file data or raw byte buffers. It cannot leak document contents or read private pixels.
- Deterministic Reproducibility: Because the output of the planner is an explicit JSON object, the user can inspect, modify, or replay the pipeline deterministically without re-querying the AI model.
- Near-Zero Latency: Token generation is limited to a concise 200-to-400 token JSON structure, returning in roughly 400ms to 700ms.
System Prompt Engineering & JSON Schema Enforcement§
Feeding an LLM a catalog of 80+ tool definitions without causing context bloat, sluggish latency, or hallucinated arguments requires careful prompt engineering and schema compression.
In NexaBatch, we utilize Gemini 3.6 Flash configured with structured JSON mode (responseMimeType: 'application/json') and a low temperature (0.1). This forces the model to produce strictly parseable JSON without conversational pleasantries, introductory prose, or markdown code fence blocks.
Compressed Tool Schema Registry§
Rather than transmitting verbose OpenAPI specifications or hundreds of lines of JSON schema documentation for every tool, NexaBatch maintains an ultra-compact schema dictionary. Each tool is described by its unique ID, a single-sentence capability summary, an array of accepted input MIME types, its output MIME type, and its configurable option flags:
// Example tool definitions from the NexaBatch Registry
const TOOL_DEFINITIONS = [
{
id: "exif-strip",
desc: "Strips all EXIF/IPTC/XMP metadata from images",
inputs: ["image/jpeg", "image/png", "image/webp"],
output: "input", // Preserves input format
options: {
preserveOrientation: { type: "boolean", default: true }
}
},
{
id: "image-resize",
desc: "Resizes image dimensions preserving aspect ratio",
inputs: ["image/*"],
output: "input",
options: {
maxWidth: { type: "number", required: false },
maxHeight: { type: "number", required: false },
maintainAspectRatio: { type: "boolean", default: true }
}
},
{
id: "image-convert",
desc: "Converts image format to webp, png, jpeg, or avif",
inputs: ["image/*"],
output: "specified",
options: {
format: { type: "string", enum: ["webp", "png", "jpeg", "avif"], required: true },
quality: { type: "number", min: 0.1, max: 1.0, default: 0.85 }
}
},
{
id: "create-zip",
desc: "Packages multiple files into a compressed .zip archive",
inputs: ["*/*"],
output: "application/zip",
options: {
filename: { type: "string", default: "archive.zip" }
}
}
];
The System Prompt Architecture§
The system prompt explicitly commands the model to act as a pure compiler that converts natural language into an executable array of tool calls:
const SYSTEM_PROMPT = `You are NexaBatch, an expert compiler that converts natural language file processing requests into a linear execution pipeline.
You have access to the following tool specifications:
${JSON.stringify(TOOL_DEFINITIONS, null, 2)}
STRICT OPERATIONAL RULES:
1. Return ONLY a valid JSON array of pipeline steps. Do NOT wrap in markdown \`\`\`json blocks.
2. Each item in the array MUST have this exact schema:
{
"tool": "<valid_tool_id>",
"label": "<human-readable description of what this step does>",
"options": { "<option_key>": <option_value> }
}
3. The first step receives the user's uploaded files. Each subsequent step receives the output files of the previous step.
4. VALIDATE MIME COMPATIBILITY: Do NOT pipe the output of an incompatible tool into a subsequent tool (e.g. do not pass application/pdf into an image-only tool unless preceded by a pdf-to-image conversion step).
5. If the user's request is impossible given the available tools or files, return an array containing a single error object:
[{ "error": "Clear explanation of why the requested pipeline cannot be fulfilled" }]
`;
By enforcing responseMimeType: 'application/json', the browser receives an immediate ArrayBuffer stream that parses directly with JSON.parse(responseText), eliminating regex stripping or syntax extraction headaches.
Dynamic Dependency Loading & Just-In-Time Execution§
One of the greatest technical hurdles in client-side web development is bundle size. A comprehensive suite featuring 80+ file tools relies on a wide array of specialized client-side engines and WebAssembly binaries:
pdf-lib(Document object manipulation): ~380 KBJSZip(Client-side DEFLATE archive compression): ~100 KBTesseract.js(Optical Character Recognition): ~2.5 MB core + trained language modelsFFmpeg.wasm(Audio/video multiplexing and transcoding): ~25 MB WASM binaryPapaParse(Fast CSV streaming parser): ~45 KB
If NexaTools bundled all these libraries into the initial page load of /utility-tools/nexabatch.html, users would be forced to download over 40MB of JavaScript and WebAssembly before performing a simple image conversion.
The Just-In-Time (JIT) Dependency Resolver§
To solve this, NexaBatch employs a dynamic dependency injection mechanism. Tools declare their external CDN dependencies inside their adapter metadata. The executor analyzes the compiled execution plan and dynamically injects only the scripts required for the specific tools in the current pipeline:
class DependencyManager {
constructor() {
this.loadedScripts = new Set();
this.pendingPromises = new Map();
}
async loadDependency(url) {
// If script is already loaded and ready, resolve immediately
if (this.loadedScripts.has(url)) {
return Promise.resolve();
}
// If script is currently being fetched, join existing promise to prevent duplicate requests
if (this.pendingPromises.has(url)) {
return this.pendingPromises.get(url);
}
const loadPromise = new Promise((resolve, reject) => {
// Check if script tag exists in DOM already
const existing = document.querySelector(`script[src="${url}"]`);
if (existing) {
this.loadedScripts.add(url);
return resolve();
}
const script = document.createElement('script');
script.src = url;
script.async = true;
script.crossOrigin = 'anonymous';
script.onload = () => {
this.loadedScripts.add(url);
this.pendingPromises.delete(url);
resolve();
};
script.onerror = () => {
this.pendingPromises.delete(url);
reject(new Error(`[NexaBatch] Failed to load dependency: ${url}`));
};
document.head.appendChild(script);
});
this.pendingPromises.set(url, loadPromise);
return loadPromise;
}
async ensureDependencies(urls = []) {
await Promise.all(urls.map(url => this.loadDependency(url)));
}
}
When a user executes an image-only pipeline, the browser loads zero PDF or FFmpeg libraries. The initial payload of NexaBatch remains under 65 KB, delivering instantaneous page rendering. When a user requests a PDF merge, pdf-lib is fetched in parallel from high-speed Cloudflare edge CDN nodes in under 150ms and cached by the browser's HTTP cache for all subsequent operations.
Handling Chained Memory & Zero-Copy Blob Pipelines§
In traditional Node.js or Python backend pipelines, batch processing relies on temporary files written to the local filesystem (e.g., /tmp/step1_out.png). Operating systems use disk paging, kernel file descriptors, and OS buffer caches to move data between processes.
In a client web browser, JavaScript execution is constrained within a single tab's V8 isolated process. Modern 64-bit browsers typically impose an upper heap memory limit of roughly 1.5 GB to 4 GB per tab. If a batch pipeline attempts to hold dozens of 40MB raw uncompressed pixel arrays in memory simultaneously, the browser tab will crash with an out-of-memory Aw, Snap! error.
NexaBatch overcomes this limitation using Zero-Copy Blob Pipelines combined with aggressive garbage collection hooks.
┌────────────────────────────────────────────────────────────────────────────┐
│ ZERO-COPY IN-MEMORY BLOB PIPELINE │
└────────────────────────────────────────────────────────────────────────────┘
Step 1: Input Files (Disk / Drop Reference)
│
▼
[EXIF Stripper] ──> Emits: Blob { size: 4.1MB, type: 'image/jpeg' }
│
├────────────────────────────────┐
▼ │
[Image Resizer] ──> Reads Blob via createImageBitmap() │
Draws to OffscreenCanvas ▼
Emits: Blob { size: 1.2MB } URL.revokeObjectURL()
│ (Releases 4.1MB RAM)
├────────────────────────────────┐
▼ │
[WebP Converter] ─> Encodes via canvas.toBlob() │
Emits: Blob { size: 380KB } ▼
│ URL.revokeObjectURL()
▼ (Releases 1.2MB RAM)
[ZIP Packager] ───> Streams Blobs into JSZip
Emits final archive Blob { size: 18MB }
The In-Memory Streaming Contract§
Every NexaBatch adapter takes an array of File or Blob objects and returns a new array of File or Blob objects. Blobs are opaque, immutable file-like references maintained by the browser's native C++ storage layer outside the JavaScript heap.
// Base pattern for memory-safe image transformation
async function processImageStep(inputBlobs, targetFormat, quality) {
const outputBlobs = [];
for (const blob of inputBlobs) {
// 1. Decode image bitmap using hardware-accelerated worker thread
const bitmap = await createImageBitmap(blob);
// 2. Allocate an OffscreenCanvas matching source dimensions
const canvas = new OffscreenCanvas(bitmap.width, bitmap.height);
const ctx = canvas.getContext('2d');
ctx.drawImage(bitmap, 0, 0);
// 3. Immediately close the bitmap to free GPU texture memory
bitmap.close();
// 4. Convert canvas pixels to destination Blob format
const outputBlob = await canvas.convertToBlob({
type: targetFormat,
quality: quality
});
// 5. Wrap with metadata preserving original filename
const newName = replaceExtension(blob.name || 'image', targetFormat);
outputBlobs.push(new File([outputBlob], newName, { type: targetFormat }));
}
return outputBlobs;
}
Critical Memory Optimizations§
createImageBitmap()overHTMLImageElement: Standardnew Image()elements force the browser to decode and retain pixel buffers on the main DOM thread.createImageBitmap()decodes images asynchronously in background worker threads and provides an explicit.close()method that immediately releases texture RAM from the GPU.OffscreenCanvas: Standard<canvas>elements belong to the DOM tree and cause paint reflows.OffscreenCanvasworks completely detached from the DOM, allowing rendering operations to execute smoothly without freezing the user interface.- Explicit Pointer Revocation: After Step 2 consumes the output of Step 1, NexaBatch invalidates intermediate
Blobreferences and invokesURL.revokeObjectURL()on any object URLs created for UI previews. This signals the V8 garbage collector to reclaim physical RAM before the next step begins.
Error Recovery, Schema Validation & Fallback Routing§
AI planning is inherently non-deterministic. What happens when a user enters an nonsensical request like "Turn this MP3 voice memo into a transparent PNG image"? Or what if a network interruption causes the Gemini API to drop connection?
A production-grade web application cannot simply crash or display a cryptic stack trace. NexaBatch implements a multi-tier defense architecture:
[User Prompt Submitted]
│
▼
[Gemini 3.6 Flash Planner]
│
┌───────────────┴───────────────┐
▼ ▼
[Valid JSON Output] [API / Schema Error]
│ │
▼ ▼
[MIME Type Validator] [Fallback UI Mode]
/ \ │
(Passed) (Failed) │
│ │ │
▼ ▼ ▼
[Execute Plan] [Flag Incompatible] [Open Manual Pipeline Drawer]
1. Semantic Type Validation Gate§
Before executing any plan returned by Gemini, the client executor executes a static type-check over the proposed DAG. It verifies that:
- Every tool ID in the pipeline exists in the
NexaBatchRegistry. - Required tool arguments are present and within valid ranges (e.g.,
qualitybetween0.1and1.0). - The output MIME type of step $N$ is accepted by step $N+1$.
If a type mismatch is detected (for example, the model mistakenly attempted to pass an audio/mpeg file into an image/webp converter), the executor rejects the automated run and flags the offending step with a clear badge in the UI: "Step 2 requires image input, but Step 1 produces audio".
2. The Manual Pipeline Fallback Drawer§
If the AI planning engine fails or if the user is working offline without an active internet connection, NexaBatch provides a seamless Manual Pipeline Fallback.
Users can bypass the AI planner entirely. The UI provides an interactive visual drawer where users can:
- Drag and reorder tool cards manually.
- Click to add tools from the categorized tool picker (Media, PDF, Data, Crypto).
- Adjust numeric sliders, dropdown selectors, and text inputs with real-time validation.
- Save customized multi-step pipelines as reusable browser templates stored in
localStorage.
This hybrid approach ensures that power users maintain total control over their workflows while enjoying the convenience of natural language automation when desired.
Frequently Asked Questions (FAQ)§
How does Gemini 3.6 Flash plan multi-step file pipelines?§
Gemini 3.6 Flash acts as a semantic compiler. It is supplied with a lightweight catalog of over 80 tool schemas defining input MIME types, output types, and parameter constraints. When you submit a prompt and drop your files, the model evaluates your intent, maps required operations to the catalog, and outputs a strictly formatted JSON array defining the execution graph.
What happens if an intermediate tool step fails?§
If a tool encounters a corrupted file or an unhandled processing error during execution, NexaBatch halts the pipeline immediately and isolates the failed file. The UI displays an informative error report identifying the exact step and file that failed. Any files that completed successfully up to that point remain accessible in browser memory, preventing wasted processing time.
How does NexaBatch avoid downloading 50MB of libraries on page load?§
NexaBatch uses a dynamic dependency loader. Tool adapters declare their external library requirements (such as pdf-lib or JSZip) as metadata. The engine only loads scripts from high-speed CDNs when a specific tool is included in the active execution pipeline. Image-only workflows never download document or video libraries, keeping the initial page payload under 65 KB.
Can I manually edit or reorder the AI-generated pipeline steps?§
Yes. Before clicking Execute Pipeline, NexaBatch displays an interactive visual preview of every planned step. You can modify any option (such as compression quality, resize dimensions, or watermark text), remove unnecessary steps, drag cards to reorder the execution flow, or add new tools directly from the manual drawer.
What are the file size and memory limits for browser execution?§
Because processing occurs inside the browser tab's memory sandbox, operations are governed by available system RAM and browser tab heap limits (typically 2 GB to 4 GB). NexaBatch streams files through memory pipelines and aggressively frees intermediate Blobs, enabling smooth processing of hundreds of photos or multi-megabyte documents without tab crashes.
Does NexaBatch work without an internet connection?§
The execution engine is 100% client-side and can run offline once tool scripts are cached. However, the initial AI planning phase requires an internet connection to communicate with the Gemini 3.6 Flash API. If you are offline, you can use the Manual Pipeline Fallback to build, configure, and execute multi-tool workflows without any network connectivity.
Conclusion & The Future of In-Browser AI Agents§
The architecture of NexaBatch illustrates a compelling vision for the future of web applications. For over a decade, the standard engineering response to complex computational workflows was to spin up expensive serverless functions, container clusters, and cloud storage buckets.
By combining low-latency frontier AI models like Gemini 3.6 Flash for control-plane reasoning with in-browser WebAssembly, WebCodecs, and HTML5 APIs for data-plane execution, we can deliver experiences that are orders of magnitude faster, completely free to operate, and cryptographically private by design.
Experience the planner in action today on NexaBatch. If you are interested in exploring our related developer utilities, check out the JSON Formatter & Validator, the LLM Token Counter, and our hash-generator.