• 8 mins read
  • Published

PewDiePie's Ajax Tests the Case for Local AI

Noel Sharkey Technology, AI and robotics editor Science.Report

Post by Noel Sharkey

PewDiePie's Ajax Tests the Case for Local AI Science.Report © science.report
PewDiePie's Ajax Tests the Case for Local AI © science.report

PewDiePie has announced Ajax, a locally run assistant based on Alibaba's Qwen3.5-9B, after describing two OpenAI account suspensions during development activity he characterized as distillation.

Felix Kjellberg, better known as PewDiePie, says OpenAI suspended his account twice while he was developing Ajax, a small language model intended to run locally as an always-on assistant. One suspension was reportedly linked to distillation, the practice of using outputs from a larger model to help train a smaller one. No public statement from OpenAI confirming the specific incident was identified in the available reporting, so the account remains Kjellberg's characterization rather than a verified OpenAI case history.

Ajax is not yet a finished commercial assistant. Its website was still presenting the project as coming soon on October 2, and the available reports did not identify a public model download, confirmed hardware requirements, independent benchmark results, or a confirmed license for the weights. There is also no published study describing a test population, task sample size, confidence intervals, p-values, or laboratory affiliation for Ajax's performance.

  • A local model with narrow ambitions

    Ajax is being built for Odysseus, Kjellberg's self-hosted AI workspace. It is described as being based on Alibaba's Qwen3.5-9B model and intended to interact with tools for web searches, browsing, email, calendars, and other routine tasks. The design assumes that a smaller model can be useful when it is connected to software tools rather than asked to perform every task inside one enormous cloud system.

    That is a credible engineering distinction. Parameter count is not a direct measure of reliability, and a model designed for a particular software environment may need less capacity than a general-purpose frontier system. A nine-billion-parameter model may also be easier to run locally than a much larger model, but the parameter figure alone does not establish memory use, response latency, energy consumption, accuracy, or operational safety.

    The project's central claim is therefore about deployment architecture rather than demonstrated intelligence. Processing is intended to remain on the user's hardware instead of sending everyday activity entirely to a cloud provider. Local execution can reduce dependence on remote services and may limit routine data transfer, but it does not automatically guarantee privacy: logs, browser permissions, plug-ins, telemetry, and connected accounts can still expose sensitive information.

    Research groups such as Stanford's Center for Research on Foundation Models distinguish between a model's raw capabilities and the behavior produced by a complete system of prompts, tools, permissions, and evaluation procedures. That distinction is especially important for Ajax because its proposed usefulness depends on external actions. A browser agent that can search, send mail, or modify a calendar must be assessed not only for language quality but also for authorization boundaries, error recovery, and resistance to prompt injection.

    No independent testing of those properties has been reported. The available material does not establish how often Ajax completes tasks correctly, how frequently it takes an unsafe action, how much human supervision is required, or whether performance changes across hardware configurations. Until those measurements are published, claims about reliability remain prospective rather than empirical.

  • The distillation dispute

    Kjellberg says he attempted to use outputs from a larger AI model as seed material for Ajax. In machine learning, knowledge distillation transfers selected behavioral patterns from a larger teacher model into a smaller student model. The process can reduce computational requirements, but it does not guarantee that the student will reproduce the teacher's factual accuracy, reasoning ability, refusal behavior, or calibration.

    Distillation quality depends on the prompts used to generate training examples, the diversity and filtering of those examples, the student model's training objective, and the evaluation protocol. A rigorous assessment would normally define held-out tasks before testing, report the number and composition of examples, compare against suitable baselines, and disclose uncertainty or failure rates. None of those details has been publicly reported for Ajax.

    According to the available reporting, Kjellberg displayed an OpenAI email that attributed one account deactivation to activity involving distillation. He says the account was restored after an appeal and then suspended again after he used the model to generate what he characterized as seed data for Ajax. The chronology rests on Kjellberg's account and the email shown in his launch video, not on a public technical explanation from OpenAI.

    The dispute exposes a practical limit on the idea that open model development can freely build on closed systems. Distillation is a recognized technical method, yet the contractual status of generated outputs depends on how they were obtained and what the resulting model is intended to do. A technique can be technically ordinary while its use remains restricted by a provider's terms of service.

  • Refusals and unfinished testing

    Kjellberg is also modifying Ajax's refusal behavior. The project describes those restrictions as ablated, while reporting indicates that he used the open-source Heretic tool to perform an automated process known as abliteration. The stated aim is a less restricted assistant, although Kjellberg says boundaries remain around requests involving harm to other people or oneself and that legal advice rules out dangerous actionable instructions.

    Those assurances are not a safety evaluation. Removing or weakening refusal behaviors can alter responses in ways that are difficult to predict from ordinary conversations. A meaningful safety study would need adversarial prompts, benign controls, repeated trials, clear harm categories, independent review, and transparent reporting of both false refusals and dangerous failures. No such dataset, trial count, statistical analysis, or peer-reviewed result is currently available for Ajax.

    Safety researchers at MIT and reporting standards associated with journals such as Nature emphasize the importance of testing systems under realistic conditions rather than relying on demonstrations alone. For an agent with access to email, calendars, and web tools, evaluation would also need to measure permission handling, susceptibility to indirect instructions embedded in web pages, and the consequences of partially completed tasks.

    The project plan includes additional reinforcement learning followed by another round of decensoring, quantization, and benchmarking. These are future development steps rather than completed results. Quantization can reduce the memory footprint of a model by representing weights with lower numerical precision, but its effects on accuracy and tool use must be measured for the specific model and hardware. Reinforcement learning can change behavior, yet the direction and consistency of those changes depend on the reward design and training data.

    The numerical description is limited but important: Ajax is based on Qwen3.5-9B, a model with around nine billion parameters. That is substantially smaller than the largest systems used by major AI companies, but the figure says nothing by itself about Ajax's accuracy, latency, energy use, tool reliability, or privacy. No independent score, trial count, failure rate, or hardware specification was available when the project was announced.

    Local deployment is a worthwhile engineering direction because it can place computation and data closer to the user, but it introduces its own maintenance and security responsibilities. Users may need to manage model updates, software dependencies, access permissions, corrupted outputs, and hardware failures. A locally stored model can reduce exposure to a cloud operator without eliminating risks from malware, compromised integrations, or unauthorized access to the device.

    Ajax's most meaningful test will therefore be repeatable performance on defined tasks with disclosed failures, rather than the fluency of a launch demonstration. A useful public evaluation would report latency and resource consumption across specified hardware, tool-call success rates, rates of unsafe or unauthorized actions, refusal consistency, and comparisons with both the unmodified Qwen model and relevant cloud assistants.

    As of October 3, 2026, the available material still describes Ajax as an announced and demonstrated project rather than a confirmed public release with downloadable weights. The weights, license, system requirements, independent benchmarks, and evaluation data remain unavailable in the reporting reviewed here. Ajax currently demonstrates a recognizable development idea and a sharp conflict over model outputs; it does not yet demonstrate a trustworthy everyday assistant.

    The broader lesson is measured rather than revolutionary. Smaller self-hosted systems may offer useful combinations of lower deployment scale, tool integration, and user control. But privacy, reliability, and safety are properties of the entire deployed system, not automatic consequences of local execution or a reduced parameter count. Ajax makes those questions visible while leaving the decisive empirical answers for a future release and independent evaluation.

  • Related articles