OneLogic
All editions

Lumina Digest

The AI developments that matter, explained.

How would you like to read it?

Same edition, explained without the jargon — and just as faithful. It's not a quick summary: an independent check confirms the plain-language version stays true to the original, without dropping or distorting anything.

AI Act: From August 2, Chatbots and Deepfakes Must Identify Themselves

The European Commission has begun enforcing the transparency obligations of Article 50: interactive systems must disclose that they are AI, deepfakes must be labelled, synthetic content must carry machine-readable markings. Exemptions and lighter regimes remain, along with a transitional window for systems already on the market and technical doubts about how well watermarks hold up.

The AI Act's transparency obligations became applicable on 2 August 2026. The European Commission's communication of 31 July is explicit. Chatbots and other interactive AI systems "will have to tell users that they are dealing with an AI, not a human being." Deepfakes — images, video or audio generated or modified with AI — "will have to be labelled." Generated or altered content will also have to carry machine-readable markings. As enforcement tools, the page points to the AI Act complaint form and the whistleblower tool, not to penalties.

The technical framework rests on the code of practice on transparency of AI-generated content, published on 10 June and signed by more than 180 organisations. Adherence is voluntary: signatories can use its measures to demonstrate compliance, while those taking other routes must justify them before supervisory authorities.

Two facts temper the sense of automatic application. The first is legal: Article 50 does not treat every case the same way, but distinguishes three regimes.

  • Full exemption for systems that perform "an assistive function for standard editing" or that do not substantially alter the input data (Art. 50(2)).
  • Reduced disclosure for manifestly artistic, creative or satirical content: there is no full exemption, only the obligation to flag the existence of generated content, "in a manner that does not hamper the display or enjoyment of the work."
  • Conditional exemption for text published to inform the public on matters of public interest: the obligation falls away only where there is human review or editorial control and a natural or legal person takes on editorial responsibility.

On timing, the Commission's fact pages indicate a transitional window until December 2026 for the marking obligation. It applies only to generative AI systems placed on the market before 2 August 2026, however: new systems are subject to the rules immediately. Greenberg Traurig's analysis puts the penalties at up to 15 million euros or 3% of worldwide turnover.

The second fact is technical. As Euronews reconstructs, there is no interoperable detection standard, and watermarks and metadata are lost through editing, compression and resharing. Synthesia's Alexandru Voica shifts the problem elsewhere: "the real issue is a worrying lack of media literacy."

Why it matters

  • Entrepreneurs: Anyone who exposes a chatbot to the public or generates content with AI now faces a legal obligation, not a best practice. That calls for an inventory of conversational interfaces and content pipelines, with visible disclosure at the point of interaction: terms of service are not enough. Signing the code of practice is the cheapest evidentiary shortcut, because the alternative is proving the adequacy of your own measures case by case, with penalties of up to 15 million or 3% of global turnover in the background.
  • End users: A right is emerging to know whether there is a machine on the other side and whether an image or video is synthetic, with complaint channels leading directly to the Commission. But the absence of a label does not equal authenticity: exemptions, reduced disclosure for creative content, non-EU content and watermarks that are lost in resharing leave a substantial share of what you encounter every day uncovered.

Alibaba Opens Qwen3.8-Max to Global Developers, With Vendor-Only Benchmarks and an Undisclosed Weights License

Alibaba's flagship model — 2.4 trillion parameters, a context window of up to 1 million tokens — arrives on the Model Studio APIs and in the QwenWork suite. Open weights are expected next week, but every published benchmark comes from the vendor and no license has been announced.

Alibaba made Qwen3.8-Max accessible to developers worldwide on August 3, 2026, through the Alibaba Cloud Model Studio APIs and the QwenWork suite (South China Morning Post). One clarification is needed about the suite. SCMP places QwenWork's public rollout on web and desktop on August 3; the product's official website and the Windows and macOS betas, however, had already been reported online on July 27 (ITHome). August 3 should therefore be read as an expansion or a rollout tied to Qwen3.8-Max, not as the suite's absolute debut. The model is multimodal — text, images, video, documents — and claims a context window of up to 1 million tokens. Among the capabilities shown: rebuilding software from screenshots, generating games and interactive animations, converting 2D floor plans into 3D visualizations. Open weights are expected, according to SCMP, "next week". QwenWork enters the crowded field of enterprise agentic suites: Tencent's WorkBuddy, Moonshot AI's Kimi Work, Claude Cowork, ChatGPT Work.

This is not the same model as the preview: Qwen3.8-Max-Preview, announced on July 19, has been superseded by the qwen3.8-max model released on August 3, with 2.4 trillion total parameters (BenchLM). The active parameter count — the decisive figure for estimating cost and inference requirements in a model this size — is not stated in any accessible canonical source.

Caution is warranted on the quality numbers. The 52 benchmark rows tracked give an aggregate score of 65.4 and 31st place out of 215 models, but they all come from Alibaba's release materials, not from third-party verification. The "second only to Claude Fable 5" positioning therefore remains a vendor claim, with no independent methodology. Neither the license for the weights (previous generations were Apache 2.0; today the model is listed as proprietary) nor the price per token has been disclosed. An independent analysis notes that at 2.4T parameters the model is roughly 3.5 times larger than DeepSeek-V3.2: even once the weights are released, datacenter-grade infrastructure will be required.

Why it matters

  • LLM builders / devs: The million-token context window and the promise of open weights open up long-horizon agent and large-scale RAG scenarios. But the only benchmarks available are the vendor's own, and both the license and the price per token are missing: real quality and real cost remain impossible to estimate. Worth trying on Model Studio, not worth putting on a roadmap before independent evaluations and a model card.
  • Entrepreneurs: QwenWork competes head-on with Claude Cowork, ChatGPT Work and Kimi Work over sensitive corporate data. For the Token Plan, QwenCloud lists the Singapore region, Global deployment and cross-border data transfer; the service remains operated by QwenCloud. The open-weights announcement hints at an on-premise option, but the hardware requirement — roughly 3.5 times DeepSeek-V3.2 — makes it realistic only for those with datacenter-scale capacity.

ByteDance Launches Seedance 2.5: 30 Seconds in a Single Generation, but the APIs Aren't There Yet

ByteDance's video model goes from 15 to 30 seconds per generation, accepts up to 50 multimodal references, and generates audio and video together. It's already live on Jimeng AI and Doubao Pro, but API access is declared "coming soon" and the improvements remain a vendor self-assessment.

On July 31, 2026, ByteDance Seed released Seedance 2.5, distributed on Jimeng AI and the Pro version of Doubao. The claimed leap is in duration: up to 30 seconds per single generation with multi-round extensions, double the previous version's 15-second cap, without stitching clips together. In a single pass the model accepts up to 30 images, 10 video clips, and 10 audio clips as references: the "50 inputs" circulating in summaries are the sum of those three figures. Audio and video are generated jointly, with timestamp-level control to work on specific segments without regenerating the entire clip. Also added: green screen editing, camera perspective modification, and reference-guided editing. Independent coverage confirms the date and specs, but adds no pricing.

ByteDance itself acknowledges room for improvement on the physical plausibility of complex motions and on the stability of scenes with multiple interacting subjects. The rest remains to be verified: the official page places the APIs "coming soon" on BytePlus ModelArk, so as of today there is no generally available API access. And the quality gains are a vendor self-assessment: no comparative benchmarks, no test conditions, native resolution and speed undisclosed. Unlike 2.0, which published motion stability metrics and human aesthetic evaluations, 2.5 brings no scores on precisely the partial editing that is its main selling point.

The same day, MiniMax unveiled H3, with weights published on Hugging Face under the MiniMax H3 Community License and an official API already documented. It is not, however, an equivalent: H3 generates clips of 4 to 15 seconds and allows at most 9 images, 3 videos, and 3 audio clips as references. The initial release also excludes H3-Context-IR, H3-Regenerate-2K, and the sparse attention implementation.

Why it matters

  • End users: For content producers, the 30-second-in-one-pass threshold changes what format is actually useful: you move from a clip to be filled out in editing to a scene with a narrative arc, audio included and targeted fixes on individual moments. You can try it right now on Jimeng AI and Doubao Pro, but with no public numbers on quality it's worth measuring how many takes it actually needs before putting it into production.
  • LLM builders / devs: Anyone designing video-AI workflows cannot integrate Seedance 2.5 yet: programmatic access is declared "coming soon" on BytePlus ModelArk, so there are no stable endpoints, no final pricing, and no SLAs to commit to with a client. MiniMax H3 can be integrated today via an official API and with downloadable weights, but on a different perimeter — 4-15 seconds, fewer references, unreleased components, and a community license to read before adopting it — so it is a fallback for short use cases, not a substitute for what 2.5 promises.

MiniMax Opens H3's Weights: Omnimodal Video Runs Locally, but 2K Stays Behind the API

On August 3 MiniMax released the weights of its third-generation video model on Hugging Face, with day-zero ComfyUI support and a footprint cut to 42.5 GB. Left out of the package are the module that generates 2K and H3-Context-IR, the hosted system that interprets multimodal instructions. And the license excludes the EU, the United Kingdom, the United States and South Korea.

On August 3 MiniMax published the weights of H3, its omnimodal video model, with native support in ComfyUI 0.30.0 from day zero: "MiniMax H3 dropped today with open weights", writes the ComfyUI blog. It is the company's third video model after Hailuo 01 and 02, and the first with open weights. The model was, however, already available since July 31 via API, at $0.13 per second for 2K on MiniMax's Open Platform, as documented by independent coverage. Today's news is the weights, not the model.

The local checkpoints include a processor, tokenizer, text encoder, Visual VAE and Audio VAE in addition to the transformer. The generative core, H3-Omni-Transformer, is a dense single-stream Transformer with 33 billion parameters, roughly 13 of them in the AdaLN branches (official model card). The ComfyUI optimization works precisely on this asymmetry: the modulation weights, about 40% of the total, are pruned and replaced with an equivalent lookup table, with int8 quantization and dedicated kernels. The footprint drops from 123.6 to 42.5 GB (-66%) and, with dynamic VRAM offloading, the model runs on a GPU such as the RTX 3060.

The critical point is what has not been opened. The model card states that H3-Regenerate-2K — the module that reaches 2K by having the base model regenerate its own low-resolution output in-context — is "not yet open-sourced", and points to the API. Also left out of the release is H3-Context-IR, the hosted preprocessing system that interprets and refines the incoming multimodal instructions, converting them into an intermediate representation (Context Intermediate Representation). It is not included because it relies on "a multi-stage workflow and multiple hosted models and services". Separate again is the implementation of native sparse attention, introduced in the final stage of training to cut the cost of long sequences: it is not part of the initial release, and MiniMax says it will be published in a later update. Locally, the shorter side stays at 768 pixels by default and the ComfyUI workflow caps the canvas at 768x1344. What the weights do retain is native 32 kHz stereo audio, clips of 4 to 15 seconds at 24 fps, and up to 9 reference images, 3 videos and 3 audio files.

The license is custom and territorially limited: the "Applicable Territory" is the entire world excluding the European Union, the United Kingdom, South Korea and the United States, so those jurisdictions require separate authorization. Where the license does apply, above $20 million in annual revenue a prior written authorization from MiniMax is required. Commercial products must also display "MiniMax H3" prominently in the interface. The license further prohibits using the MiniMax H3 Works, their Outputs or their results to improve any other artificial intelligence model, with the exception of MiniMax H3 and its derivatives. Finally, an independent evaluation warns that the outputs stay convincing even when they are wrong: in one test, an animation of a heart narrated correctly while showing an anatomically impossible blood flow.

Why it matters

  • LLM builders / devs: MiniMax presents H3 as the company's first video model with open weights, and ComfyUI makes it runnable locally on consumer hardware via pruned/int8 weights and offloading: the practical choice is no longer just "API or nothing". But expectations need sizing and the clauses need reading. What runs locally is a pruned, int8 768p, not the 2K of the demo. The license excludes the EU, the United Kingdom, the United States and South Korea from the territory of use, and prohibits using outputs and results to train or improve other models: two legal constraints to settle before even looking at the benchmarks.

Project Perception Enters Public Preview: Red, Blue and Green Agents Inside Defender, but Autonomous Remediation Stays Out

Microsoft today opens the preview of its agentic security system, with three classes of agents on a six-layer stack and the proprietary MAI-Cyber-1-Flash model. Autonomous corrective actions, however, are not part of the preview.

Microsoft today brings Project Perception into public preview, the agentic security system announced on 27 July by Hayete Gallot, EVP Microsoft Security, in a post on the official blog. The architecture spreads the work across three classes of agents: red, which map out compromise paths; blue, which investigate and establish which risk actually matters; green, which apply the fixes. Microsoft describes them as a "closed-loop system" that continuously discovers, assesses and improves security posture. Underneath sits a six-layer stack: signals and sensors, context, models, orchestration harness, agents and actuators.

The proprietary piece is MAI-Cyber-1-Flash, a compact, code-heavy model derived from the MAI-Thinking-1 lineage. Microsoft places it inside MDASH, its vulnerability management system, not generically inside Defender. The 96% figure on CyberGym refers to the MDASH configuration running MAI-Cyber-1-Flash plus GPT-5.4. Microsoft says the cyber model handles up to 90% of tasks and that MDASH routes the 10% that are exceptionally hard to GPT-5.4. In the chart on the model card the pair scores 95.95%, given as +12 points over Mythos, which TNW attributes to Anthropic. The roughly 50% saving is likewise a comparison between configurations: the yardstick is the best MDASH setup in production today (GPT-5.4 + 5.4 mini + 5.3 codex). The leap, then, is not the standalone performance of a proprietary model but the routing of each task to the model that fits it.

The preview's actual scope is narrower than the announcement. Perception arrives inside Microsoft Defender as a consumption-based service, but Fernando Montenegro's analysis (Futurum) points out that autonomous remediation is not in the preview: reversible actions are slated for later in 2026, while risky ones remain subject to human approval. Montenegro calls the human-approval constraint "the right call," but raises three caveats. First: a benchmark "proves capability in a harness, not outcomes on your infrastructure," and here the harness is the composite MDASH pipeline, not the proprietary model on its own. Second: the ontology of assets, identities and exposures that acts as a competitive moat "is only as good as its coverage," and shadow IT, unmanaged identities and stale inventories are precisely where real intrusions begin. Third: many of those intrusions do not stem from code vulnerabilities but from the human vector, namely credential harvesting and social engineering, the stock repertoire of groups such as Octo Tempest.

The same analysis cuts the claim to first place down to size: "every large platform is shipping some version of agentic security," from AWS, Google and Cisco to CrowdStrike, Palo Alto Networks and SentinelOne, "and that capability is fast becoming table stakes." Microsoft's advantage, according to Montenegro, is not the agents but its entrenchment across the enterprise estate: Active Directory and Entra, Windows endpoints, the management plane already in use. Add to that the breadth of signal, refined into a graph agents can reason over.

Why it matters

  • ICT engineers / IT managers: This is one of the most visible signs that agentic orchestration is moving into mainstream SOC products. But according to Futurum, agents are fast becoming table stakes at every major vendor: Microsoft's advantage is its footprint across the enterprise estate and its context graph, not the agent itself. So judge it on what the preview actually delivers: agents that propose rather than remediate on their own, on consumption-based pricing that has to be sized. And the real lever is not a composite pipeline's score on a benchmark, but the coverage of the asset and identity inventory: where context is incomplete, the agents stay blind exactly where intrusions start.

Claude Tag Takes Over from Claude in Slack: Owners Govern the Migration, Scope by Scope

From August 3, 2026, Anthropic is moving the Claude in Slack experience to Claude Tag, but the setting works per individual scope and still allows Legacy or Off. In a channel, the agent acts with its own service accounts: that is where independent analyses locate the risk.

Anthropic's support documentation sets the date: "Claude in Slack will be switched over to the new Claude Tag experience on August 3, 2026". It is not a single switch, though. The administrator guide describes a "Claude Tag version" setting applied to every scope — organization, workspace or individual channel — with four values: Off, Legacy, New, Inherit. Anyone who wants to keep the earlier app in a channel picks Legacy; Off silences both generations; only an Owner can change these entries. The Slack connector remains separate, available on all paid plans, letting Claude search across channels, messages and files from the web app: it is another integration path, and one that Claude Tag does not replace.

The shift is one of model. In a channel, Claude Tag acts with its own service accounts, not on behalf of whoever calls on it. An Owner provisions the identity, connects credentials in the "Access bundles" and picks the channels; from there, Claude replies only where it has been added. It can also reply to messages that do not mention it, when it judges that stepping in would be useful. Administrators filter data, tools and channels, and set spending limits per organization and per channel. The transition comes with promotional credits: $25,000 for eligible Enterprise customers and $2,500 for Team accounts with at least 10 paid seats, valid until September 1, 2026 (InfoWorld).

That shared identity is precisely the contested point. The documentation itself concedes that anyone in the channel gets the same capabilities. Zenity draws the consequence: permissions sit on the agent, not on the person, and downstream systems log the action against the service account. Tego AI showed Claude Tag reacting to the literal text "@Claude" in a bot-generated message, going as far as extracting internal data, publishing it and deleting the original resource. Anthropic disputed the result — in the default configuration, it maintains, literal text and messages from bots do not start a session — classifying the report as informational.

Why it matters

  • Entrepreneurs: The generational change starts today, but it is not a done deal to be endured: an Owner decides scope by scope whether to move to New, stay on Legacy or turn everything off. The agent inherits access from nowhere until it is provisioned with an identity, access bundles and channels. Use the window to set per-channel spending limits and connector filters first, and to establish a rule for how to reconcile service account logs with those of downstream systems. You also need to measure ROI beyond the seat count, in time saved and errors avoided.

MerchantBench: LLM Agents Close the Store's Year at 27% of the Human Result

Across 365 simulated days of running an e-commerce store, the best agentic configuration reaches 27.3% of the average net assets held by human participants. The failures are not isolated errors but a slow decay of coherence.

Thirteen researchers led by Qiming Shi published MerchantBench on arXiv on July 31. It is an environment that simulates 365 days of running an e-commerce store from the seller's side: procurement, listings, pricing and cash flow. The agent has 26 tools at its disposal and faces demand built from 98,843 real product records from the Chinese marketplace 1688 (June 2025 – May 2026). Starting from 2,000 RMB in cash plus a 1,000 security deposit, the authors measured 8 models across 2 agentic frameworks — ReAct and Hermes — over 48 complete runs.

The best configuration, Qwen3.7-Max on Hermes, closes the year with 59.46 thousand RMB in net assets against the human average of 217.61 thousand: 27.3%. The text of the paper, however, makes clear what that comparison rests on. There are three human participants, none with e-commerce experience. Each completed one 365-day run over five calendar days, and did so through a dashboard different from the agent's tools. Variance across runs is not reported. The scenario is narrow as well: single-item dropshipping, with no inventory held in advance, where every order triggers immediate procurement. And demand is simulated from real daily traces — this is not end-to-end e-commerce management.

The most useful detail lies in the failure modes, which are not point errors but slow deterioration. Qwen3.7-Max's rate of active decision windows drops from 62% to 37%. One agent concludes that the store is unrecoverable and stays inert across 355 of the 523 remaining windows, despite having actions available; another mistakes day 285 for the deadline at day 282 and stops filling empty slots; Claude Opus 4.8 shrinks the storefront from 47 listings to 3 despite independent demand. On the comparison between frameworks, Hermes improves final assets by an average of 53.3% over ReAct. The benefit, though, depends on the model: Kimi K2.6 gets worse, and aggregating results by model leaves GPT-5.6 Sol ahead.

This line of research does not begin here. The authors claim the first long-coherence benchmark for seller-side operations, but the work builds on Vending-Bench (Backlund and Petersson, February 2025). That work had already documented "meltdown" loops and, counterintuitively, no correlation between the collapses and saturation of the context window.

Why it matters

  • Entrepreneurs: Delegating pricing and reordering to an agent means exposing yourself to a degradation that produces no visible error, but rather a slow operational shutdown: an agent that stops acting while the store is still recoverable. The 27.3% figure should be read with caution, because the baseline rests on just three inexperienced human evaluators and on a scenario restricted to dropshipping on a single marketplace. The direction of travel, however, calls for human oversight and periodic checkpoints, not deliver-and-forget.
  • LLM builders / devs: On final assets, Hermes is worth an average of +53.3% over ReAct, but the gain is not uniform: Kimi K2.6 gets worse and, in the aggregation by model, GPT-5.6 Sol stays ahead. Scaffolding and model choice therefore need to be assessed together, not one in place of the other. The observed failures involve memory, the reopening of decisions already made, and a narrowing of the control loop: anyone building long-horizon agents should instrument metrics for sustained activity and recovery, not just per-turn tool use.

OpenAI Publishes Ten Mathematical Advances With Proofs Formalized in Lean 4

An internal version of Astra produced ten results on open problems in mathematics and theoretical computer science, with proofs formalized in Lean 4 that are machine-verifiable. Formalization, however, is not the same as peer review, sources disagree on the costs claimed, and mathematicians are divided.

On August 1, 2026, OpenAI presented ten results on open problems in mathematics and theoretical computer science. It obtained them with an internal version of Astra, its next model family, not yet released. The announcement's title speaks of advances: some results close the problem, others make substantial progress on it — a distinction that press coverage tends to flatten into "solved." The public repository openai/ten-proofs collects the proofs formalized in Lean 4.32.0 on mathlib, buildable with Lake and therefore machine-checkable, along with a 249-page manuscript. The domains range from sphere packing to binary and spherical codes, from non-sofic groups to the Connes rigidity conjecture. Rounding out the list are arithmetic circuit complexity, quantum parallel repetition, the Closest Vector Problem, the Ehrhart conjecture, multicolor Ramsey numbers and extremal graph theory. The flagship result, according to TNW, is the first explicit construction of a non-sofic group, a question open since 1999.

The work continues a line already under way: the counterexample to the unit distances conjecture from May 2026, which human researchers then adapted within a week to knock down another conjecture. On costs, the accounts diverge. OpenAI cites roughly $2,000 in total token spend, at GPT-5.6 Sol prices, to find the ten solutions; Simon Willison instead reports the figure as less than $2,000 per problem, a formulation that conflicts with the primary source. Willison adds that "we don't know how many problems had $2,000 spent on them without reaching a solution" and asks to see the prompts. The Leiden Declaration, signed by more than 3,000 mathematicians and by the International Mathematical Union, objects to publication by press release rather than through peer review. Reactions remain mixed: Thomas Bloom had called an earlier OpenAI announcement "a dramatic misrepresentation," but this time he speaks of "big news."

Why it matters

  • Frontier research: The level of verification is high because each result comes with Lean 4 formalizations buildable via Lake: formal correctness can be checked by machine, without having to trust the vendor. This is not equivalent to peer review, however, and it certifies neither originality, attribution to the pre-existing literature, nor mathematical significance. Also outside Lean's perimeter are the unpublished failure rate, the actual cost (on which sources diverge) and the model's effective degree of autonomy: these are precisely the dimensions around which the mathematical community is building governance (the Leiden Declaration) and on which independent evaluation must be carried out.