Perceptron Mk1.5 Explained: Vision-Language Model Capabilities and Pricing

Perceptron Mk1.5 Explained: Vision-Language Model Capabilities and Pricing

Perceptron Mk1.5 Explained: Vision-Language Model Capabilities and Pricing

Most vision-language models were built to describe pictures. Perceptron Mk1.5, released on 25 September 2026 by Perceptron AI of Bellevue, Washington, was built to be a perception layer for machines that move: drones, quadruped robots and smart glasses. It accepts text, images, video and audio, and it answers not only with prose but with points, boxes, polygons, video clips and timestamped object tracks. That output contract is the interesting part, because a robot cannot act on a paragraph.

It matters now because the physical-AI conversation has shifted from “can a model caption this frame” to “can it follow my hand through forty seconds of first-person video and tell me where the screwdriver went.” Mk1.5 is a concrete, priced, API-available attempt at that, and the vendor has published a dense set of benchmark numbers.

This post separates what Perceptron has actually published from what circulates on social media, walks through the architecture of the output interface, works the pricing and latency arithmetic, and gives you a method for evaluating the model before you wire it into a perception stack.

What this covers: the confirmed facts, lineage from Mk1, the structured-output design, benchmarks and their caveats, access and cost, failure modes, a comparison framework, and an evaluation checklist.

Context and Background

Perceptron AI was founded in November 2024 by Armen Aghajanyan (co-founder and CEO) and Akshat Shrivastava (co-founder and CTO), both former research scientists at Facebook AI Research, according to the company’s May 2026 launch press release. The first model, Perceptron Mk1, launched on 12 May 2026 and was pitched as performance competitive with frontier models from Google, Anthropic, OpenAI and Qwen at a fraction of the cost. The press release named six target domains: manufacturing and industrial analytics, media and content processing, robotics and automation, geospatial analysis, security and surveillance, and device-based agentic tooling. It did not state pricing for Mk1 and did not disclose funding.

Mk1.5 narrows and sharpens that pitch. Perceptron’s own description, repeated by OpenRouter, calls it an embodied reasoning model for physical agents. “Embodied” here means the model is expected to reason about a scene from the viewpoint of an agent that is inside it, typically a camera mounted on a moving body, rather than from the neutral viewpoint of a curated photo or a third-person clip.

The broader market context is the vision-language model (VLM) category. General-purpose frontier models such as Gemini and GPT accept video, and open-weight families such as Qwen-VL and Molmo have pushed grounding and pointing into the open. The gap Perceptron is targeting is the gap between “understands video” and “emits machine-usable geometry over time.” If you are new to how these systems fuse modalities, our explainer on multimodal AI architecture for vision, language and audio fusion covers the building blocks, and our survey of physical AI and vision-language-action models for robotics explains where a perception model sits relative to a policy model that outputs motor commands.

A note on sourcing. The primary source for this article is Perceptron’s own launch post, “Introducing Perceptron Mk1.5”, and the OpenRouter model page, both read on the day of writing. Almost every benchmark number below comes from the vendor’s own post. Independent replication was not available when this was written, and I flag that wherever it matters. For background on the long-running benchmarks mentioned, see the Ego4D project, which defined much of the first-person video research agenda.

What is verified and what is not

Item Status Source
Release date 25 Sep 2026 Verified Vendor post, OpenRouter
Model ID perceptron-mk1.5 Verified Vendor post
Input modalities: text, image, video, audio Verified Vendor post, OpenRouter
Output: text, points, boxes, polygons, clips, object tracks Verified Vendor post
Price $0.15 per million input, $1.50 per million output tokens Verified Vendor post, OpenRouter
Context window Vendor says 32K; OpenRouter lists 36,864 Discrepancy, see below
Max output tokens 8,192 Verified on OpenRouter only OpenRouter
Parameter count Not stated by vendor; a social post claims 35B Unverified
Architecture, base model, training data Not disclosed Vendor post silent
Open weights None announced Not found

The distinction matters because a parameter count repeated on social media is not a specification. The vendor’s post, as read, does not state one, so treat the “35B” figure as an unconfirmed claim until Perceptron or an independent party confirms it.

The Core Design: A Perception Model With a Geometric Output Contract

Perceptron Mk1.5 is a multimodal reasoning model that takes text, image, video and audio and returns text plus optional structured geometry, namely points, boxes, polygons, clips and timestamped object tracks, so that downstream robot or application code can consume its answers directly instead of parsing prose.

Perceptron Mk1.5 request and response flow from multimodal inputs to structured geometry

Figure 1: Perceptron Mk1.5 as a perception layer. Multimodal inputs go in, a reasoning step optionally calls tools or sub-agents, and the answer comes back as text plus machine-readable geometry.

Figure 1 shows the data flow as published. Inputs on the left are the four modalities. The model reasons over them within a single multimodal context, optionally calls tools declared in OpenAI JSON Schema format, optionally dispatches sub-agents, and returns text plus annotations. The inside of the “reasoning” box is a black box: Perceptron has not published the architecture, so the figure shows interfaces only, not layers.

Why structured geometry changes the integration problem

A conventional VLM returns a sentence such as “the person on the left is holding a red cup.” A planning module then needs a second system, usually an open-vocabulary detector, to turn “red cup” into pixels. That two-model chain doubles latency, doubles failure surface, and creates a grounding mismatch: the language model and the detector may disagree about which object the phrase refers to.

When the same model that understood the question also emits the box, the referent is by construction the one it reasoned about. OpenRouter’s description confirms annotations are returned inline when requested through specific request parameters, alongside function calling and JSON Schema structured outputs. For an engineer this means the response is a typed object, not free text, and a validator can reject malformed geometry before it reaches a controller.

The practical consequence is a cleaner contract at the system boundary. A drone autopilot or a glasses overlay renderer needs coordinates and timestamps. A model that natively speaks coordinates removes a translation layer, and every translation layer is a place where bugs and latency accumulate.

Tracks, not per-frame detections

The most distinctive claim in the vendor post is that Mk1.5 “emits object tracks as timestamped geometries, rather than one-off, per-frame detections.” The difference is substantial. A per-frame detector gives you independent boxes at each frame and leaves identity association to a separate tracker such as a Kalman-filter pipeline or a re-identification network. A track output asserts that the box at t=3.2 s and the box at t=7.9 s are the same entity.

This is why the reported tracking benchmarks are the headline. The post reports an F1 of 0.647 on Molmo2-Track and a HOTA of 0.628 on Ref-DAVIS17, and says the model outperforms competitors on ReasonVOS. HOTA (Higher Order Tracking Accuracy) is a standard multi-object tracking metric that balances detection accuracy and association accuracy, so a good HOTA means the identities held together, not merely that boxes landed on objects. The tracking is also referring-expression driven: you describe the target in language (“the person in the blue jacket who picked up the box”) and the model follows it.

Reasoning, tool use and sub-agents

Beyond perception, the post says Mk1.5 supports tool calls declared in the OpenAI JSON Schema format, “without retraining” for arbitrary tool functions. It also states that the model “can deploy sub-agents to parallelize and accelerate task completion.” The post offers a video example but does not describe the mechanism. My reading, labelled as interpretation, is that a coordinating instance spawns additional model calls over sub-tasks (for example, segmenting a long video into windows) and merges the results, which would explain part of the reported latency gains on long inputs. That is inference, not a published design.

The tool-use number worth noting is on the web-search benchmarks. The vendor reports 56.0 percent accuracy on LiveVQA-W and a gain of 36.1 points on MMSearch with tools versus without. A 36-point swing says that, for knowledge-heavy visual questions, access to search matters more than any perceptual improvement. Note that a third-party catalogue lists the OpenRouter endpoint as not supporting built-in web search, which implies the search tool must be supplied by your own application through function calling rather than being a hosted feature.

Deeper Analysis: Benchmarks, Latency and Cost Worked Through

What the vendor reports, in one place

All figures in the table are as published by Perceptron in its launch post. None has been independently reproduced as of writing, and the post names its comparison set (several Gemini, Qwen, GPT, AV-Flamingo and Molmo variants) without this article being able to verify each comparison row. Read the table as a map of where the vendor chose to compete, not as a leaderboard.

Capability area Benchmark Reported Mk1.5 result
Object tracking in video Molmo2-Track F1 0.647
Referring video tracking Ref-DAVIS17 HOTA 0.628
Egocentric hand localization Perceptron_Ego hand_box 0.9433
Egocentric video QA EgoSchema 80.4 percent
Egocentric video QA, hard subset EgoSchema_hard 63.75 percent
Audio-visual DailyOmni 74.675 percent
Audio-visual WorldSense 50.325 percent
Audio-visual OmniBench 51.058 percent
Audio-visual hallucination AVHBench 0.56
Web-grounded visual QA LiveVQA-W 56.0 percent
Video reasoning (standard) Perception Test 67.6 percent
Video reasoning (standard) MVBench 73.4 percent
Video reasoning (standard) TempCompass 78.0 percent
Video reasoning (hard subsets) NExT-QA 64.2 percent
Video reasoning (hard subsets) VSI-Bench 85.4 percent

Two entries deserve a warning. The vendor post lists EgoSchema at 80.4 percent in one place and a hard-subset figure of 61.3 percent in another section, alongside a separate EgoSchema_hard of 63.75 percent. These are different slices or evaluation settings, and the post does not make the distinction obvious in the text I could extract. Do not quote a single “EgoSchema score” for this model without citing which row you mean.

The second warning is about the hand-localization claim. The vendor says Mk1.5 is “50 percent better than the strongest Gemini model” on hand localization. That is a relative improvement on a benchmark, Perceptron_Ego, that appears to be in-house, so it cannot be compared with published third-party results. It is plausible and interesting, but it is a vendor-defined test.

Reading the audio-visual numbers honestly

Audio-visual understanding is where the vendor is most candid. The post states: “We do not claim to be at the frontier of audio-visual understanding in this first release, but our current public capabilities are already strong enough to be useful.” Scores near 50 percent on WorldSense and OmniBench, both multiple-choice style evaluations with several answer options, show a model that works but is not dominant. If your application depends on reasoning about sound synchronized with video, such as diagnosing a machine by its noise while watching it, test that path specifically instead of assuming the headline video numbers transfer.

Latency: the arithmetic behind “2 to 5 times faster”

The post compares Mk1.5 with Mk1 on end-to-end request completion. The reported pairs are below, with the speedup computed from them.

Task Mk1 Mk1.5 Speedup computed
Chat 5.2 s 1.1 s 4.7x (5.2 / 1.1 = 4.73)
Image question answering 1.7 s 0.51 s 3.3x (1.7 / 0.51 = 3.33)
60-second video 19 s 9.1 s 2.1x (19 / 9.1 = 2.09)

The vendor’s multipliers match the raw seconds, which is a useful sanity check. The pattern is informative: the largest gain is on short text-heavy turns, and the smallest is on video. That is what you would expect if the speedup comes from model and serving improvements that matter most when there is little input to encode, while video requests are dominated by the cost of ingesting and encoding frames. A 60-second clip at 9.1 seconds is roughly 6.6 times faster than real time (60 / 9.1), which is fine for offline analytics and for “review what just happened” features, but it is not a closed-loop control rate.

That last point is critical for robotics readers. A 1.1 second chat turn or a 0.51 second image answer is far too slow for a stabilizing control loop that runs at tens or hundreds of hertz. Mk1.5 belongs in the deliberative layer, the part of a robot stack that decides what to look for and where, with fast onboard controllers handling the actual motion. The vendor’s claim that it operates across drones, quadrupeds, smart glasses and phones “without platform-specific retraining” is a statement about perception transfer, not about running inside the control loop. The post does not disclose the hardware, region or load conditions behind the latency figures, so treat them as indicative.

Latency comparison of Perceptron Mk1 versus Mk1.5 across chat, image Q and A, and 60-second video tasks

Figure 2: Reported end-to-end latency for Mk1 and Mk1.5 by task type, as published by the vendor. The relative gain shrinks as input volume grows.

Pricing: a worked example

Mk1.5 is priced at $0.15 per million input tokens and $1.50 per million output tokens, identically in the vendor post and on OpenRouter. Output tokens cost ten times input tokens, so output length is the lever. Perceptron does not publish how many tokens a second of video or an image consumes, and I could not verify a tokenization rate, so the example below uses clearly hypothetical numbers.

Suppose a single analysis request consumes 20,000 input tokens (video plus prompt) and returns 600 output tokens of text and annotations. Input costs 20,000 / 1,000,000 x $0.15 = $0.003. Output costs 600 / 1,000,000 x $1.50 = $0.0009. The request totals $0.0039, about four tenths of a cent. At 100,000 such requests a month the bill is $390. The figures are illustrative; the point is the structure. With a 32K-token context window as the ceiling, one request can never carry more than about 32,768 input and reasoning tokens, so the maximum input cost per call is bounded near $0.005 regardless of how long your source video is. Cost per request is cheap; the real cost driver is how many windows you must send to cover long footage.

Structured annotations are not free either. A polygon with dozens of vertices, or a track with a box per sampled timestamp, consumes output tokens at the higher rate. A tracking request over a long clip can therefore be output-heavy. Budget for the output side when you request dense geometry, and prefer boxes or points over polygons when a coarse location is enough.

The context window discrepancy

The vendor post lists 32K tokens of multimodal context. OpenRouter lists a context length of 36,864 tokens and a maximum output of 8,192. The difference is 4,096 tokens, and 36,864 equals 32,768 plus 4,096, which suggests the router counts input window plus an output allowance. That is my inference, not a documented fact. In practice, plan against the vendor’s 32K for input and check the live endpoint metadata before relying on the larger number.

Thirty-two thousand tokens is modest next to the million-token windows advertised by some frontier models. For video, it forces a windowing strategy: split a long recording into segments, query each, and merge results, which is presumably where the “sub-agents” feature is meant to help. If your workload is hour-long footage, the cost of coordination and cross-window identity continuity is on you unless you use the vendor’s agentic features.

Windowing strategy for long video under a 32K context limit with sub-agent fan-out and merge

Figure 3: A windowing pattern for footage longer than the context allows. Segments are processed in parallel, tracks are stitched across boundaries, and a final reasoning pass resolves identity conflicts. This is a recommended application pattern, not a documented Perceptron feature.

Access and Deployment

Perceptron describes Mk1.5 as available through the Perceptron Platform at platform.perceptron.inc and an updated software development kit, with a Python package named perceptron on PyPI and a browser demo at perceptron.inc/demo. Enterprise deployment goes through the vendor’s sales contact. It is also live on OpenRouter, which per its model page forwards requests directly to Perceptron as the single provider, so there is no multi-provider routing or fallback behind that endpoint.

I found no announcement of open weights, a license for self-hosting, quantization options or on-device builds. The vendor does mention smartphones and smart glasses as deployment platforms, but the evidence in the post is that these devices use the model, presumably by calling a hosted API, not that the model runs on them. If your threat model requires on-premises inference, disconnected operation, or data residency guarantees, those are questions to put to the vendor directly, because nothing public answers them.

A third-party catalogue also lists the OpenRouter endpoint as lacking prompt caching and code execution. That matters for cost: without caching, a repeated long system prompt or repeated reference frames are billed at the full input rate every call.

Perceptron Mk1.5 deployment paths through the platform API SDK and OpenRouter

Figure 4: Access paths. Applications reach the same hosted model through the Perceptron Platform API and Python SDK or via OpenRouter; no self-hosted path has been announced.

Trade-offs, Gotchas, and What Goes Wrong

The evidence is vendor-authored

The largest limitation is epistemic. Nearly every number in this article originates with Perceptron. Vendor benchmark selection is not neutral: teams choose evaluations where their design shines, sometimes build in-house sets such as Perceptron_Ego, and choose comparison baselines and prompts. That does not make the numbers wrong. It means the correct posture is “promising, unreplicated,” and the right response is to run your own evaluation on your own footage.

Contamination and benchmark transfer

Public video benchmarks such as EgoSchema, MVBench and NExT-QA have been circulating for years, and their questions and clips can leak into training corpora. The vendor post says its training aims to make models “learn efficiently and adapt” rather than memorize knowledge, but it publishes no data card and no contamination analysis. High scores on a public benchmark are weak evidence that the model will track the specific, cluttered, motion-blurred footage from your factory floor.

First-person video is adversarial in specific ways

Egocentric footage has properties that break assumptions. Hands constantly occlude objects. Cameras shake and rotate. Objects leave the frame and return minutes later. Lighting changes as a wearer moves between rooms. A tracking model can score well on curated sequences and still lose identity after a long occlusion. When you test, deliberately include leave-and-return events, near-identical objects (two similar tools on a bench), and low light.

Latency, not accuracy, may be the blocker

As worked out above, even the improved 0.51 second image answer cannot sit in a fast control loop. If your drone must avoid an obstacle in 100 milliseconds, the model cannot be the avoidance system. Treat it as the planner that says “land near the marked pad” while onboard code does the flying. Also remember network round-trips: every call to a hosted API adds transit latency and a failure mode when connectivity drops, which for mobile robots is routine.

Hallucinated geometry

A model that outputs coordinates can be confidently wrong in a new way. A fabricated box looks authoritative because it is numeric. The AVHBench result of 0.56 hints that audio-visual hallucination is not solved, and the vendor admits it is not at the frontier there. Always validate geometry against basic sanity checks: boxes inside image bounds, plausible size, temporal continuity, and agreement with a cheap independent detector for safety-relevant decisions.

Safety, privacy and surveillance exposure

The press release lists security and surveillance among target domains, and smart glasses capture bystanders who never consented. Sending first-person video of people to a third-party API raises data protection duties under regimes such as GDPR and sector rules such as HIPAA. Perceptron has not published, in anything I could read, retention terms, regional processing options or a model-level safety evaluation. Ask for them in writing before shipping.

Cost surprises

Low per-token prices hide volume effects. Continuous monitoring multiplies request counts; dense polygon output multiplies the expensive output tokens; and the lack of prompt caching on at least one catalogue’s description of the endpoint means repeated context is rebilled. Model the monthly cost on realistic traffic, not on a single request.

How It Compares and How to Evaluate It

Because the architecture and parameter count are undisclosed, comparison has to be made on interface and behavior, not internals. The matrix below is a decision aid built from what each class of model is publicly known for, with Mk1.5 rows grounded in the vendor facts above. It is deliberately qualitative; I am not asserting rankings the sources do not support.

Use case Perceptron Mk1.5 General frontier multimodal model Open-weight VLM you host
Referring object tracking in video, geometry out Native tracks, a core design goal Usually text answers; geometry support varies Possible with grounding-capable models, more integration work
Long video, more than 32K tokens of context Needs windowing or sub-agents Some offer much larger windows Depends on model and your GPU memory
On-premises or disconnected operation No announced self-host path Usually hosted only Yes, you control the deployment
Cost predictability at volume Low listed token prices, no caching listed Varies widely Fixed hardware cost, your ops burden

Where a frontier general model likely wins is breadth: open-ended world knowledge, long documents, complex multi-step reasoning in text. Where an open-weight model wins is control and data locality. Where Mk1.5 aims to win is the middle: low-cost, low-latency, geometry-native perception for agents, with the quality claims still awaiting independent confirmation.

A seven-step evaluation protocol

First, assemble a golden set of 100 to 300 clips from your own environment, labelled with the geometry you need. Second, define the metric that matches your decision: HOTA or identity switches for tracking, mean intersection over union for boxes, and accuracy for question answering. Third, run Mk1.5 and at least one baseline on identical prompts and sampling settings. Fourth, slice results by condition: occlusion, low light, motion blur, crowd density. Fifth, measure latency at the 50th and 95th percentile from your deployment region, not from the vendor’s figures. Sixth, record cost per successful task, not per request, since retries inflate real cost. Seventh, repeat after a vendor model update, because hosted models change under you.

Practical Recommendations

Start by deciding which layer of your system the model occupies. If it is the deliberative perception layer, which answers “where is the thing I was told to find, and is it the same thing as ten seconds ago,” then Mk1.5’s design fits. If you need reflex-speed control, keep that on the device. For wearable and spatial computing products, the same split applies, as we discuss in our look at spatial computing on the Apple Vision Pro for manufacturing, where on-device sensing handles tracking and cloud models handle understanding.

Treat launch claims with the discipline we recommend for any frontier release. Our piece on fact-checking viral frontier-model claims lists the habits: find the primary source, separate vendor-reported from independently replicated, and look for what is not said. Apply them here. The “35B parameters” figure circulating on social media is a good test case: no primary source states it, so do not repeat it as fact.

Use the structured output interface properly. Request the narrowest geometry that serves your decision, validate every annotation against a schema and sanity checks, and log the raw responses so you can audit failures. Pin the model ID, and rerun your evaluation set whenever the vendor ships an update.

Checklist before production:

  • Build a labelled golden set from your own footage and score Mk1.5 against it.
  • Measure p50 and p95 latency from your deployment region.
  • Confirm data retention, regional processing and training-use terms in writing.
  • Add geometry sanity checks and a fallback detector for safety-relevant actions.
  • Design a windowing and identity-stitching strategy for footage beyond 32K tokens.
  • Cap output tokens and prefer boxes or points over dense polygons.
  • Model monthly cost on real request volume, including retries.
  • Plan for vendor lock-in: wrap the API behind your own interface.

Frequently Asked Questions

What is Perceptron Mk1.5?

Perceptron Mk1.5 is an embodied reasoning model from Perceptron AI, released on 25 September 2026. It takes text, image, video and audio input and returns text plus structured annotations such as points, boxes, polygons, clips and object tracks. It is aimed at physical agents including drones, quadruped robots and smart glasses, and it is accessed through the Perceptron Platform, a Python SDK and OpenRouter rather than as downloadable weights.

How much does Perceptron Mk1.5 cost?

The published price is $0.15 per million input tokens and $1.50 per million output tokens, shown both in the vendor’s launch post and on OpenRouter. Output is ten times the input rate, so long structured outputs such as dense polygons cost more than the same volume of input. A hypothetical request with 20,000 input and 600 output tokens would cost about $0.0039, though real token counts per video second are not published.

What is the context window of Perceptron Mk1.5?

The vendor describes 32K tokens of multimodal context. OpenRouter lists 36,864 tokens with a maximum output of 8,192 tokens. The gap of 4,096 tokens is probably an output allowance, though that is an inference rather than a documented fact. Plan around 32K for input, and for long videos split the footage into windows and merge results in your own application or with the vendor’s sub-agent feature.

Is Perceptron Mk1.5 open source?

No open-weight release has been announced as of this writing. Access is through a hosted API on the Perceptron Platform and OpenRouter, with a Python package on PyPI for the client SDK. The vendor mentions commercial licensing and enterprise deployment through its sales team, so on-premises use may be negotiable, but nothing public specifies terms, hardware requirements or quantization options.

How many parameters does Perceptron Mk1.5 have?

Perceptron has not stated a parameter count in the launch post I read, and it has not disclosed the base model or architecture. A social media post has claimed 35 billion parameters, but that is an unverified third-party claim. Until the company or an independent analysis confirms it, treat the size as unknown and judge the model on measured behavior with your own data.

Is Perceptron Mk1.5 fast enough for real-time robot control?

Not for low-level control. The vendor reports 0.51 seconds for image question answering and 1.1 seconds for a chat turn, and 9.1 seconds for a 60-second video. Those are fine for planning and perception queries but too slow for loops that need responses in tens of milliseconds. Use it as a deliberative layer above fast onboard controllers, and measure latency from your own network location.

Further Reading

By Riju — about

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *