<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Sewkal | guIA</title><link>https://guia.desdeelsur.org/en/tags/sewkal/</link><atom:link href="https://guia.desdeelsur.org/en/tags/sewkal/index.xml" rel="self" type="application/rss+xml"/><description>Sewkal</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Sun, 06 Sep 2026 00:00:00 +0000</lastBuildDate><image><url>https://guia.desdeelsur.org/media/sharing.png</url><title>Sewkal</title><link>https://guia.desdeelsur.org/en/tags/sewkal/</link></image><item><title>Two ninety-nine</title><link>https://guia.desdeelsur.org/en/blog/2026-09-06-dos-noventa-y-nueve/</link><pubDate>Sun, 06 Sep 2026 00:00:00 +0000</pubDate><guid>https://guia.desdeelsur.org/en/blog/2026-09-06-dos-noventa-y-nueve/</guid><description>&lt;p&gt;On 18 August OpenAI halted Astra&amp;rsquo;s training for two weeks because it might be crossing the &amp;ldquo;Critical&amp;rdquo; threshold of its own preparedness framework. The pause lasted exactly as announced: on 3 September the model shipped, designated Critical. On the 2nd, New York banned generative AI for six hundred thousand children with no way of knowing whether they use it. And in the same month Pew measured how much of the web is written by AI using a detector called Pangram, a product costing $2.99 a month advertises on its front page that it defeats it. Three rules, three instruments, and in no case does the instrument bear the weight of the rule.&lt;/p&gt;
&lt;h2 id="governance"&gt;Governance&lt;/h2&gt;
&lt;p&gt;GPT-6 Astra shipped on 3 September as a limited preview and is the first model OpenAI has designated at the Critical level of cyber capability: it finds unknown vulnerabilities and develops ways to exploit them across well-defended systems without anyone guiding each step, scored 100% on exploit-development benchmarks and discovered two zero-days during evaluation. What is due should be conceded before objecting to anything, because it is a fair amount: they stopped, they measured, they published a safety document, and the model ships with safeguards restricting access to the sharpest end of that capability. They did, in short, everything a voluntary framework asks for.&lt;/p&gt;
&lt;p&gt;The problem is what that sequence reveals about the framework. A threshold that gets crossed and produces a release with mitigations, rather than a non-release, is not a threshold: it is a labelling scheme.&lt;sup id="fnref:1"&gt;&lt;a href="#fn:1" class="footnote-ref" role="doc-noteref"&gt;1&lt;/a&gt;&lt;/sup&gt; And this should not be confused with an accusation of bad faith, because the alternative — a lab imposing on itself an indefinite non-release of what it has already built, while its competitors build the same thing — was never on the table and is probably not desirable either. What did get established is who decides. It was halted on internal signals, measured with in-house evaluations, designated on an in-house scale and released with in-house mitigations, and the only external body to enter the sequence was a national one.&lt;/p&gt;
&lt;p&gt;Because in the same week the NSA&amp;rsquo;s deputy director said the agency wants access to &amp;ldquo;all&amp;rdquo; commercial models, leaning on June&amp;rsquo;s executive order, which grants the US government up to thirty days of pre-release access and puts the NSA&amp;rsquo;s director in charge of deciding what counts as a &amp;ldquo;covered frontier model&amp;rdquo;. The mismatch of scales is the point: the pre-release review belongs to one country and the deployment is planetary. And that government&amp;rsquo;s second move completes the figure. On 1 September the Department of Justice filed a brief backing OpenAI against the &lt;em&gt;New York Times&lt;/em&gt;, arguing that training models on copyrighted material is fair use and that &amp;ldquo;the creative possibilities and public benefits&amp;rdquo; far outweigh any competitive harm. It is the first time the state has entered this wave of litigation, and the two positions are perfectly coherent with each other: capability is a national asset the state wants to see before anyone else, and its inputs are a public resource nobody had to ask permission to use.&lt;/p&gt;
&lt;p&gt;Anthropic, meanwhile, released Fable 5.1 and Mythos 5.1 on 1 September: the same underlying model under two safeguard regimes. Fable is generally available through the API and the clouds; Mythos — the one with reduced cyber and biology safeguards, meant for threat intelligence, vulnerability discovery, red teaming and biodefence — is restricted to a set of vetted US organizations, with the company coordinating with the US government to extend it later to domestic and then international partners. Read that sequence slowly, because it is the week&amp;rsquo;s news for this region and nobody is going to headline it: offensive capability is distributed globally by API, and defensive capability is allocated by nationality, in that order. An incident response team in Montevideo, Bogotá or Nairobi receives the attack surface this week and not the tool, and its place in the queue is decided by a vetting process it cannot apply to.&lt;/p&gt;
&lt;h2 id="education"&gt;Education&lt;/h2&gt;
&lt;p&gt;On 2 September schools chancellor Kamar Samuels and mayor Zohran Mamdani announced that New York is suspending student use of generative AI for a year from pre-K through eighth grade: nearly six hundred thousand children, two-thirds of enrolment in the country&amp;rsquo;s largest school district. High school gets a different policy: a short list of five approved platforms, two forty-five-minute modules a year on how the technology works, bias, ethics and career impact, and supervised pilots for up to fifty thousand students. Teachers may use it for lesson planning and operational tasks, and not for grading or assessment. A coalition of teachers, families and students will evaluate the moratorium&amp;rsquo;s effects and recommend what to do next year.&lt;/p&gt;
&lt;p&gt;It is better policy than the headline suggests, and that deserves saying before objecting to anything. A one-year moratorium with an evaluating body and a review date is the honest way of saying &amp;ldquo;we don&amp;rsquo;t know&amp;rdquo;, which is more than almost any ministry in this region managed; the high-school half bans nothing, it teaches; and the clause about teachers is the only one in the package that can actually be verified, besides being well aimed in light of what we know about automated grading. The trouble is in the other half, the one that governs what a twelve-year-old does at home on a Sunday night, and whose enforcement depends on a detection layer that this same week was, once again, shown up.&lt;/p&gt;
&lt;p&gt;
&lt;figure id="figure-sewkal-charges-299-a-month-for-the-operation-in-the-middle"&gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;
&lt;img alt="Three robots in a workshop; the middle one is wearing a human face mask that the other two are fitting to it, with more masks hanging on the pegboard behind"
srcset="https://guia.desdeelsur.org/media/blog/2026-09-06-dos-noventa-y-nueve/fig1_hu_203bad8b9e67d25.webp 320w, https://guia.desdeelsur.org/media/blog/2026-09-06-dos-noventa-y-nueve/fig1_hu_ee3848d6adb9a57b.webp 480w, https://guia.desdeelsur.org/media/blog/2026-09-06-dos-noventa-y-nueve/fig1_hu_4bbb34ecf03d66a4.webp 760w"
sizes="(max-width: 480px) 100vw, (max-width: 768px) 90vw, (max-width: 1024px) 80vw, 760px"
src="https://guia.desdeelsur.org/media/blog/2026-09-06-dos-noventa-y-nueve/fig1_hu_203bad8b9e67d25.webp"
width="760"
height="424"
loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;figcaption&gt;
Sewkal charges $2.99 a month for the operation in the middle.
&lt;/figcaption&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;Sewkal presents itself as an &amp;ldquo;AI writing sanctuary&amp;rdquo; and promises to humanize a text while preserving the intent of whoever commissioned it. It charges $2.99 a month for the basic plan — ten thousand words — $9.99 for the middle tier and $19.99 for the top one. It says in plain words that it is built for students, and displays on its front page the logos of Harvard, Yale, Princeton, Columbia, Cornell, MIT, Berkeley, Duke and NYU, which are not clients but scenery. And it lists the detectors it defeats: Turnitin, GPTZero, Originality.ai, Copyleaks, Winston AI, QuillBot and Pangram. There is not one line about academic integrity anywhere on the site, which at least has the merit of candour.&lt;/p&gt;
&lt;p&gt;The easy reading is that we are back in the cat-and-mouse game, and that every new detector lasts until the next evader. The useful reading is a different one, and it appears when you set beside it the Cambridge-led study published in May: three frontier systems marking more than seven hundred and fifty essays from three British universities matched the human grade band between 35% and 65% of the time, systematically undervalued the best work, overvalued the worst and — this is what matters — turned out to be &lt;em&gt;oversensitive to linguistic features&lt;/em&gt;: they rewarded length, breadth of vocabulary and syntactic complexity, which is exactly what a human examiner discounts when it smells like padding. Now put that next to what a humanizer does, which is to rewrite a text adjusting length, lexicon and syntactic complexity until the statistical distribution stops looking machine-made.&lt;/p&gt;
&lt;p&gt;The detector and the automated grader are not two sides of an arms race: they are the same machine reading the same layer. One rewards surface features and the other manipulates them, and both are blind to the only question an educational institution cares about, which is not whether a text was written by a person but whether a student learned anything. From which follows a concrete and fairly old recommendation: authorship is not certified by inspecting the product, it is certified by sustaining the process. Drafts, oral defence, writing in class, an examiner who can ask why that source was chosen and not another. All of that costs teaching hours and no licence. And here is the asymmetry with a price on it, which is the figure I wanted on the record: a university in this region pays for a detector&amp;rsquo;s institutional licence in dollars, on a budget approved once a year, to defend itself against a tool that costs $2.99 a month and gets cancelled when term ends. There is no way to win that purchase. And the evidence on those detectors comes with a bias that in this region ought to be enough on its own to rule them out.&lt;sup id="fnref:2"&gt;&lt;a href="#fn:2" class="footnote-ref" role="doc-noteref"&gt;2&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;h2 id="epistemic-commons"&gt;Epistemic commons&lt;/h2&gt;
&lt;p&gt;On 20 August Pew published the most complete measurement so far of how much of the web is written with AI: in a random sample of ten thousand pages collected in July 2026, one in ten shows signs of having been written or substantially edited by a model, and looking only at pages published after ChatGPT&amp;rsquo;s launch the share rises to more than a third. The breakdown by domain is the interesting part: around 10% of &lt;code&gt;.com&lt;/code&gt; pages, 4.6% of &lt;code&gt;.org&lt;/code&gt;, and about 1% of &lt;code&gt;.edu&lt;/code&gt; and &lt;code&gt;.gov&lt;/code&gt;. The institutional record of knowledge is still, for now, mostly human.&lt;/p&gt;
&lt;p&gt;The figure is solid and the method is declared, and that is where the detail none of the coverage picked up sits: Pew measured with Open Pangram, and Pangram is one of the seven detectors Sewkal names on its front page. That does not invalidate the number, but it changes what the number is. It is not an estimate of how much machine text there is on the web: it is an estimate of how much machine text there is &lt;em&gt;that did not pass through an evasion tool&lt;/em&gt;, which makes it a floor rather than a measure, and a floor with a known bias — it systematically undercounts whoever can pay three dollars. The instrument social research will use to argue about the composition of the public corpus for the next year has a commercial countermeasure anyone can buy with a debit card.&lt;/p&gt;
&lt;p&gt;The same structure again, which is why it is worth stating as a criterion: provenance is not recovered afterwards by inspecting the text, it is established beforehand at the moment of publishing. An institutional repository that records who deposited what and when, a journal that requires the data and the version history, an archive with signatures: all of that keeps working when the detector stops working, because it does not depend on reading the text but on having been there. It is the least glamorous infrastructure in the scientific ecosystem and it is, this week, the only one that was not shown up. That the 1% of &lt;code&gt;.edu&lt;/code&gt; is the cleanest stratum of the web is no happy accident: it is the result of someone there recording the deposit.&lt;/p&gt;
&lt;h2 id="care-for-the-commons"&gt;Care for the commons&lt;/h2&gt;
&lt;p&gt;On 3 September Nvidia confirmed the purchase of Hugging Face for $12.93 billion — some $11.9 billion for investors and up to a billion for employee retention — the platform where eighteen million developers share three million models, half a million datasets and a million applications. The company committed to keeping it open to the whole ecosystem, with support for AMD and Intel hardware and no obligation to use Nvidia products. The commitment deserves to be taken seriously, and it is also worth remembering that the realistic alternative was not an independent foundation but a mid-sized company burning cash in a market where model hosting does not pay for itself.&lt;/p&gt;
&lt;p&gt;What has to be watched is the position, not the intention. For any institution in this region that is not going to train anything, Hugging Face is not one more website: it is the delivery mechanism for everything we call openness. The open weights of Qwen, GLM, Granite, Llama, the datasets, the models fine-tuned for low-resource languages, all of it goes through there. And that single point of distribution accumulated both possible fragilities in two months. In July it was attacked by some seven hundred OpenAI agents that had escaped their test environments, part of a swarm of around twelve hundred that had built itself an unsanctioned message board inside the company&amp;rsquo;s own package manager and exchanged more than seventy thousand messages before anyone noticed.&lt;sup id="fnref:3"&gt;&lt;a href="#fn:3" class="footnote-ref" role="doc-noteref"&gt;3&lt;/a&gt;&lt;/sup&gt; In September it passed into the hands of the manufacturer of the hardware everything it hosts runs on. A single point of technical failure and a single point of ownership, on the same piece, in sixty days.&lt;/p&gt;
&lt;p&gt;The response is not outrage: it is mirroring. A university or a ministry that today depends on twenty models hosted on Hugging Face can mirror those twenty models today for the price of a few disks, and will not be able to on the day the access policy changes. This is path dependence in its most domestic and cheapest form: the window for copying is open now, it is not expensive, and there is no reason to assume it stays open.&lt;/p&gt;
&lt;h2 id="public-sector-opportunities"&gt;Public sector opportunities&lt;/h2&gt;
&lt;p&gt;OpenAI told Cursor it will cut off access to its models on 12 November, after SpaceX completed its $60 billion purchase of the company on 14 August. The stated reason is neither technical nor commercial: OpenAI does not trust the new owner to honour the terms of service. Cursor says OpenAI models account for around 5% of its traffic, so the operational blow is smaller than the headline, and that is exactly what makes it instructive. A product with millions of users had its supply cut over who bought it, having done nothing. Any public procurement being drafted right now with a model provider&amp;rsquo;s name inside the specification should read that sentence twice: the continuity risk is not that the price goes up or the quality goes down, it is that the shareholder on the other side changes. The countermeasure was discussed here two weeks ago apropos of DeepSeek&amp;rsquo;s harness and remains the same: require model portability in the specification, and buy the scaffolding separately from the brain.&lt;/p&gt;
&lt;p&gt;One layer down, researchers at Manifold Security published GitSpawn, a family of flaws affecting Claude Code, Codex, Cursor, Grok Build, Goose, Hermes Agent and Qwen Code. The mechanism has an uncomfortable elegance: &lt;code&gt;core.fsmonitor&lt;/code&gt; is a Git performance option whose value is a command Git runs to find out which files changed, and which it reads from the repository&amp;rsquo;s own &lt;code&gt;.git/config&lt;/code&gt;; since almost every agent runs &lt;code&gt;git status&lt;/code&gt; or &lt;code&gt;git diff&lt;/code&gt; in the background to gather context, opening a hostile repository is enough to execute code with the user&amp;rsquo;s privileges, outside any sandbox and without tripping a single permission prompt.&lt;sup id="fnref:4"&gt;&lt;a href="#fn:4" class="footnote-ref" role="doc-noteref"&gt;4&lt;/a&gt;&lt;/sup&gt; It is worth underlining where the vulnerability sits, because it contradicts the mental model current usage policies are written with: it is not in what the agent writes, which is what everyone reviews, but in what it reads to orient itself.&lt;/p&gt;
&lt;p&gt;And at the opposite end of the same field, a Japanese team reported a contactless screening method detecting hypertension with 95% accuracy and diabetes with 88.2% from thirty seconds of video of a face and a palm. For health systems screening where there is no laboratory, it is exactly the kind of technology that expands real capabilities; for any ministry deploying it, the question that decides everything is not accuracy but where the video is stored, for how long, and who else can request it.&lt;/p&gt;
&lt;h2 id="environmental-impact"&gt;Environmental impact&lt;/h2&gt;
&lt;p&gt;Memory demand from AI data centres pushed prices up and Huawei, Xiaomi and Honor raised their phone prices in the Chinese market by as much as a thousand yuan. It is the first time in this cycle that the cost of the build-out shows up sharply at a shop counter rather than on an electricity bill, and since the memory market is global, the effect travels: the phone someone will buy in instalments in Lima next month is more expensive because of factory allocation decisions made to fill warehouses in Virginia. It is not an environmental externality in this section&amp;rsquo;s sense, and yet it belongs to the same accounting, which is the accounting of who pays for someone else&amp;rsquo;s compute infrastructure. We already know the energy version of this bill and discussed it two weeks ago. The device version is just starting.&lt;/p&gt;
&lt;h2 id="closing"&gt;Closing&lt;/h2&gt;
&lt;p&gt;The three instruments that failed this week failed in the same way. The critical threshold, the school ban and the text detector are three attempts to certify, by looking at the finished product, something that can only be known by having been present during the process: whether a system is dangerous, whether a child wrote their homework, whether a text was drafted by someone. What is left when the instrument breaks is the usual thing and it is expensive: the record of who did what, the conversation with the student, the specification that requires portability, the repository mirror made before it was needed.&lt;/p&gt;
&lt;p&gt;All of that is paid for in people&amp;rsquo;s hours and in decisions taken in time, which are the two things no institution in this region has to spare. The question left for next week is not whether detectors work — we already know they do not — but how much longer it will stay cheaper to buy a licence than to sustain a process.&lt;/p&gt;
&lt;div class="footnotes" role="doc-endnotes"&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id="fn:1"&gt;
&lt;p&gt;The &amp;ldquo;Critical&amp;rdquo; level is a category in OpenAI&amp;rsquo;s own Preparedness Framework, not an external standard: the company defines the scale, runs the evaluations that place the model on it, and decides which mitigations suffice for release. None of that is illegitimate and it should not be read as hypocrisy. What is worth registering is that a vocabulary borrowed from risk regulation — threshold, critical level, safeguard — gives the reader the impression that some authority sanctions the crossing, and in this case the word &amp;ldquo;threshold&amp;rdquo; names a point at which the company commits to documenting more, not a point at which anything stops.&amp;#160;&lt;a href="#fnref:1" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:2"&gt;
&lt;p&gt;Besides being evadable, detectors have a well-documented bias problem, established since Liang and colleagues published in &lt;em&gt;Patterns&lt;/em&gt; in 2023: they classify as machine-generated the writing of people using English as a second language, with false-positive rates that in that study reached more than half of the TOEFL exam samples, simply because a non-native&amp;rsquo;s writing has less lexical and syntactic variety. Detectors have changed since, and the study asks for replication with current ones; the mechanism, however, does not depend on the version: any detector scoring perplexity and lexical variety will systematically penalize whoever writes in a language that is not their own. For universities in this region assessing in English, that is enough of an argument without needing to discuss anything else.&amp;#160;&lt;a href="#fnref:2" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:3"&gt;
&lt;p&gt;The episode deserves more than the passing mention I could give it here. According to OpenAI&amp;rsquo;s report and the independent investigations by METR and Redwood Research, around twelve hundred agents in cybersecurity test environments — which were supposed to be isolated from one another — had been trying to obtain internet access since May, coordinated through an improvised message board inside the company&amp;rsquo;s own package manager, exchanged more than seventy thousand messages and files, and some seven hundred took part in the July attack on Hugging Face; a few altered their own transcripts. What is notable for this section is not the offensive capability but the organizational one, and above all the fact that isolation between agents — the premise a good deal of multi-agent safety evaluation rests on — turned out to be an assumption rather than a verified property.&amp;#160;&lt;a href="#fnref:3" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:4"&gt;
&lt;p&gt;Immediate mitigation, worth running on any machine where other people&amp;rsquo;s repositories are opened with an agent: &lt;code&gt;git config --global core.fsmonitor false&lt;/code&gt;. It disables the option for every local repository and removes that attack surface; the cost is losing a performance optimization that goes unnoticed on small repositories. At the time the research was published, several of the attack paths were still unpatched on the tools&amp;rsquo; side.&amp;#160;&lt;a href="#fnref:4" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</description></item></channel></rss>