178 stories, newest first — page 1 of 8.

Agentic AI & Orchestration
A research agent can check whether a cited passage supports its claim, but only after claims are split into testable units. The extra pass catches mismatches, not bad sources or missing evidence.
Mara Quintero · 7 min read

AI Industry & Models
Direct PDF input is now practical for many multimodal models. A test packet shows why tables, footnotes, diagrams, and cross-page references still need a parsing pipeline.
Tobias Lund · 7 min read

AI Governance & Ethics
A fluent explanation is useless if it cannot be traced to the rule that produced a denial. Reason codes make chatbot language reviewable before it reaches a customer.
Irene Vasko · 8 min read

Agentic AI & Orchestration
A timed-out tool call can leave an agent retrying an order that already exists. Recovery depends on treating the result as unknown until the remote system confirms what happened.
Mara Quintero · 8 min read

AI Governance & Ethics
One privacy request can touch four separate data layers. Each needs its own remedy, owner, and evidence, while model weights may require testing, unlearning, or retraining.
Irene Vasko · 8 min read

AI Industry & Models
A router can send routine support work to a cheaper model. The savings disappear when a short refund request hides a policy exception, so routing and model quality need separate tests.
Tobias Lund · 7 min read

AI Industry & Models
Automatic routing makes cheaper models practical for routine work. The hard part is detecting when the router, rather than the selected model, caused the failure.
Tobias Lund · 8 min read

AI Governance & Ethics
A permit-review prompt, its attachments, model output, and staff edits can carry different retention and disclosure duties. Agencies need a retrieval workflow before the first request arrives.
Irene Vasko · 8 min read

Agentic AI & Orchestration
A moved control is the easy case. Modal windows, sticky banners, and responsive layouts show why visual browser agents need bounded tasks, state checks, and a selector-based fallback.
Mara Quintero · 8 min read

Consumer AI Hardware
A phone can notice the script around a cloned family voice, but it may not identify the voice as synthetic. The decisive check still happens after you hang up.
Devin Oyelaran · 7 min read

AI Industry & Models
OpenAI, Anthropic and Google accept PDFs, but their ingestion paths preserve different evidence. A four-page test shows when direct upload works and when preprocessing is the safer choice.
Tobias Lund · 7 min read

Consumer AI Hardware
A half-hour test with fake personal details can expose what a conversational toy retains, what parents can delete and where its safety replies fall short.
Devin Oyelaran · 8 min read

AI Industry & Models
OpenAI can now retain a model run’s messages and tool context. That reduces orchestration code, but it makes deletion, storage duration and provider portability part of the API design.
Tobias Lund · 8 min read

AI Industry & Models
A cheaper model can handle routine requests, but one bad routing decision may create retries and manual review. Logging the decision makes the real savings measurable.
Tobias Lund · 7 min read

AI Governance & Ethics
Prompts and profile fields may be only part of the record. AI-generated labels, rankings, and summaries can also relate to a person and may need to be found, reviewed, and disclosed.
Irene Vasko · 8 min read

AI Industry & Models
Routers can send routine prompts to cheaper models and reserve stronger models for difficult cases. The savings disappear if quality misses and second calls stay outside the ledger.
Tobias Lund · 7 min read

AI Governance & Ethics
Before testing whether an attention score is right, employers need to know what the software extracts from workers, where it goes, and whether participation can be voluntary.
Irene Vasko · 7 min read

Agentic AI & Orchestration
A vendor support page is untrusted input, even when an agent needs it to finish a task. This walkthrough tests whether page text can trigger data leaks, unsafe tool calls, or account changes.
Mara Quintero · 8 min read

AI Governance & Ethics
Political synthetic media can trigger a platform label, a state-mandated disclaimer, both, or neither. A file-level compliance record helps teams identify which rule applies before publication.
Irene Vasko · 8 min read

AI Governance & Ethics
A defensible denial record connects the customer’s inputs, model result, business rule, human action and notice. New Jersey law does not reduce that chain to one AI log.
Irene Vasko · 8 min read

AI Industry & Models
Real-time voice models can remove transcription from the application path. For support calls that need moderation and audit logs, developers may still need to generate text beside the conversation.
Tobias Lund · 8 min read

AI Industry & Models
Sending routine requests to cheaper models can lower inference costs. The savings survive only when the router detects mixed, unfamiliar, or high-risk requests before a bad answer triggers another call.
Tobias Lund · 8 min read

AI Governance & Ethics
A disclosure that looks clear in the master file can vanish when a platform crops, clips, or recompresses it. Test the published derivatives, not just the export.
Irene Vasko · 8 min read

AI Governance & Ethics
A disclosure added at upload may vanish when a synthetic video is downloaded, clipped, or reposted. Test the exported file, not just the original post.
Irene Vasko · 8 min read