<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Kaldryn: Blog, Research and EU Regulation Briefings</title>
    <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/</link>
    <atom:link href="https://delightful-grass-0e0338903.7.azurestaticapps.net/feed.xml" rel="self" type="application/rss+xml" />
    <description>Private, self-hosted AI from Sweden. Writing on sovereign AI, on-premise deployment and EU compliance.</description>
    <language>en</language>
    <lastBuildDate>Tue, 28 Jul 2026 17:48:15 GMT</lastBuildDate>
    <item>
      <title>Local AI in Sweden and the EU: The 2026 Guide to Sovereign Deployment</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/research/local-ai-sweden-eu-guide/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/research/local-ai-sweden-eu-guide/</guid>
      <pubDate>Sat, 04 Jul 2026 09:00:00 GMT</pubDate>
      <category>Research &amp; white papers</category>
      <category>Sovereign AI</category>
      <description>Why Swedish and European organisations are moving AI on-premise in 2026, the regulation driving it, who is deploying local AI, and how to choose a sovereign platform.</description>
      <content:encoded><![CDATA[<p>In most European boardrooms the AI question has changed shape. Two years ago it was “should we use AI?”; in 2026 it is “where does it run, and who can see the data?” Local AI, also called on-premise or sovereign AI, is the answer a growing share of Swedish and EU organisations are giving: the models run on hardware the organisation owns and controls, inside its own walls, instead of in a vendor's cloud.</p>
<p>This guide covers why that shift is happening now, which sectors in Sweden and the wider EU are actually deploying local AI, what the deployment options look like, and what to demand from any vendor who says the word “sovereign”.</p>
<h2>Why European organisations are bringing AI in-house</h2>
<p>The baseline is the GDPR. The moment an employee pastes a customer record, a patient note or a personnel file into a cloud assistant, the organisation is processing personal data on someone else's infrastructure, which triggers processor agreements, transfer analyses and, when things go wrong, fines that can reach 4% of global turnover. None of that paperwork exists for a model that never leaves the building.</p>
<p>The second driver is the transfer problem the Schrems II judgment created in 2020 and no amount of contractual engineering has fully solved since. Data residency is not data sovereignty: a US-owned cloud region in Frankfurt or Stockholm is still subject to US law, including the CLOUD Act's reach into data held abroad by American providers. For Swedish public bodies and German regulated industries alike, that residual risk has quietly shaped cloud policy for half a decade.</p>
<p>The third driver has a timetable: the EU AI Act entered into force in August 2024 and its obligations arrive in phases. Self-hosting does not exempt anyone from the Act, but it makes the compliance file dramatically simpler, because the deployer controls the logs, the model versions and the data flows it must document.</p>
<table><thead><tr><th>Date</th><th>What starts applying</th></tr></thead><tbody><tr><td>2 Feb 2025</td><td>Bans on prohibited AI practices; AI-literacy duties</td></tr><tr><td>2 Aug 2025</td><td>Obligations for general-purpose AI models; governance rules</td></tr><tr><td>2 Aug 2026</td><td>Most high-risk system obligations (Annex III)</td></tr><tr><td>2 Aug 2027</td><td>High-risk rules for AI embedded in regulated products</td></tr></tbody></table>
<p>Around the AI Act sits a thicket of sector rules pushing the same direction: NIS2 for critical infrastructure, DORA for financial institutions since January 2025, professional privilege for law firms, and secrecy legislation for the public sector. In Germany, works councils and BSI security expectations add their own gravity toward on-premise deployment.</p>
<div data-callout><p>None of this makes cloud AI illegal. It makes cloud AI expensive to justify: every cloud assistant now carries a standing compliance burden, while a local deployment removes most of the questions instead of answering them.</p></div>
<h2>Who is deploying local AI in Sweden and the EU</h2>
<p>Adoption follows a simple pattern: the sectors deploying local AI first are the ones that were never able to use cloud AI properly in the first place.</p>
<ul><li>Healthcare, clinics and care providers running assistants over patient documentation, where GDPR Article 9 special-category data rules out external processing.</li><li>Legal, firms that cannot put privileged client material into any third-party service, but still want drafting, summarisation and research over their own matter files.</li><li>Financial services, banks and insurers balancing DORA, banking secrecy and outsourcing rules, where an on-premise model sidesteps the vendor-risk file entirely.</li><li>Public sector, Swedish agencies and municipalities whose cloud caution since Schrems II is well documented, and who increasingly specify sovereignty in procurement.</li><li>Manufacturing and defence, companies protecting product IP and export-controlled information, often on air-gapped networks where cloud AI is simply impossible.</li></ul>
<p>What these organisations run locally is no longer a compromise: private ChatGPT-style chat for employees, retrieval-augmented answers over internal documents, drafting and summarisation, and OpenAI-compatible APIs that let internal tools call a local model exactly as they would call a cloud one.</p>
<p>The vendor landscape splits into three camps. Hyperscalers now market “sovereign cloud” regions, better than nothing, but the stack, the keys and the legal entity questions remain debated. European API providers keep inference on EU soil but still off your premises. And a third camp puts the AI inside your building: Kaldryn, built by Pontén Solutions in Stockholm, is a Swedish example of that approach, shipping Kaldryn One, a locker-sized appliance that runs a full ChatGPT-class assistant, 200+ open models and an OpenAI-compatible API entirely inside the customer's network.</p>
<h2>Appliance or DIY: the two ways to go local</h2>
<p>Every organisation that decides to run AI locally faces the same fork. The DIY route assembles open-source parts, an inference server such as vLLM or Ollama, open-weight models, a vector database, a chat UI, and it works, as every lab prototype proves. The hidden cost is everything around the model: SSO, DLP, audit logging, model provenance, GPU drivers, upgrades and an on-call rota. It is a realistic path only for organisations with a platform team to spare.</p>
<p>The appliance route buys the same outcome as a product: pre-built GPU hardware with the platform pre-installed, deployed in minutes and owned like any other piece of IT equipment.</p>
<ul><li>DIY: maximum flexibility; you own integration, hardening and lifecycle. Budget for engineering time, not licenses.</li><li>Appliance: turnkey, plug in power and ethernet, index your documents, serve the whole team. Kaldryn One is built for exactly this.</li><li>Hybrid: start with one appliance for a department; scale to racks or your own servers later without changing the platform.</li></ul>
<p>The economics favour local more than most buyers expect. Per-seat cloud AI compounds: Microsoft 365 Copilot lists at $30 per user per month, roughly $10,800 a year for a 30-person company before add-ons. A local deployment inverts the model: a fixed hardware cost, then electricity, with unlimited users and unlimited prompts. In the Nordics, where power is comparatively cheap and largely fossil-free, running your own GPUs is also the low-carbon option.</p>
<h2>A buyer's checklist for sovereign AI</h2>
<p>Whatever vendor you evaluate, including us, these are the questions that separate sovereignty from sovereignty-washing:</p>
<ol><li>European entity: is the vendor an EU/EEA company, contracting under a European legal system?</li><li>GDPR Article 28: is a data-processing agreement offered, or better, is the architecture such that the vendor never processes your data at all?</li><li>Training clause: does the contract state, in writing, that your data is never used to train anyone's models?</li><li>The unplug test: ask the vendor to disconnect the internet during the demo. Sovereign AI keeps working.</li><li>Model provenance: are model weights signed and verified, so you know exactly what is running?</li><li>Ownership and exit: do you own the hardware, and does the system keep running if the vendor disappears?</li><li>EU AI Act artefacts: does the platform produce the logs and technical documentation the Act expects from deployers?</li><li>Security integration: SSO/SAML, DLP, immutable audit logs and SIEM export, sovereignty includes your security team.</li></ol>
<h2>Sweden, Germany and the Nordic moment</h2>
<p>There is a reason so much of this movement is centred on northern Europe. Sweden combines very high digitalisation with a strong privacy culture and a public sector that writes sovereignty into procurement. Germany brings the Mittelstand's instinct to keep IP in the house, works councils with a real say over employee-data tooling, and BSI-shaped security expectations. Add Nordic electricity prices and the region's fossil-free grid, and the calculus lands in the same place: for a Swedish or German organisation in 2026, the question is rarely whether local AI is possible, it is which box to plug in.</p>
<h2>The bottom line</h2>
<p>Cloud AI will remain the default for workloads that carry no sensitive data. But for the European organisations that handle patient records, client files, financial data or state secrets, the direction of travel is set: AI is becoming something you own, not something you rent. Sovereignty stopped being a philosophy the day it became a procurement checkbox.</p>
<div data-callout><p>Kaldryn publishes reference architectures and compliance checklists for self-hosted AI in this research hub, and Kaldryn One is the fastest way to see local AI running on your own desk: power, ethernet, ten minutes.</p></div>]]></content:encoded>
    </item>
    <item>
      <title>The AI Act deadlines just moved: what the 2026 Digital Omnibus changed</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/ai-act-deadlines-moved-2026-digital-omnibus/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/ai-act-deadlines-moved-2026-digital-omnibus/</guid>
      <pubDate>Thu, 02 Jul 2026 09:00:00 GMT</pubDate>
      <category>EU regulations</category>
      <category>EU AI Act</category>
      <description>In mid-2026 the EU pushed back the AI Act's high-risk deadlines. Here is what moved, what did not, and what deployers should actually do about it.</description>
      <content:encoded><![CDATA[<p>If you built a 2026 compliance plan around the original AI Act calendar, some of it is now out of date. In late June 2026 the Council gave its <a href="https://www.consilium.europa.eu/en/press/press-releases/2026/06/29/artificial-intelligence-council-gives-final-green-light-to-simplify-and-streamline-rules/">final green light</a> to a package the EU has been calling the Digital Omnibus, and one of the things it does is push back the hardest deadlines in the <a href="https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng">AI Act</a>. Worth knowing exactly what shifted before you rewrite anything.</p>
<h2>What moved</h2>
<p>The two deadlines that moved are the ones for high-risk systems, which were always going to be the heaviest lift.</p>
<ul>
<li>The rules for standalone high-risk systems listed in Annex III (think biometrics, critical infrastructure, hiring, credit scoring, access to essential services) were due to apply from 2 August 2026. They now apply from 2 December 2027.</li>
<li>The rules for high-risk AI embedded in regulated products under Annex I (medical devices, machinery and the like) moved from 2 August 2027 to 2 August 2028.</li>
</ul>
<p>The main reason is prosaic. The harmonised technical standards that companies need in order to actually prove conformity were running late at CEN and CENELEC, the European standards bodies. You cannot ask firms to meet a bar that has not been drawn yet. So the co-legislators drew the bar's deadline forward in time instead. The European Parliament's own <a href="https://www.europarl.europa.eu/news/en/press-room/20260316IPR38219/meps-support-postponement-of-certain-rules-on-artificial-intelligence">write-up</a> of the March committee vote lays out the logic.</p>
<h2>What did not move</h2>
<p>This is the part people miss. Plenty of the AI Act is already in force and none of it was delayed.</p>
<p>The ban on prohibited practices has applied since 2 February 2025. The AI literacy duty in Article 4, which says staff who use AI systems need a baseline understanding of them, has applied since the same date. The obligations on general-purpose AI models kicked in on 2 August 2025, along with the governance and penalty framework. And the transparency duties in Article 50, the ones that say you have to tell people when they are talking to a machine and label AI-generated content, still apply from 2 August 2026. They were not touched.</p>
<div data-callout><p>So if your reaction to the delay is "good, we have more time," check which bucket your use actually falls into. An internal assistant that drafts and summarises is usually not a high-risk Annex III system at all. Which means the deadlines that moved were never your deadlines. The literacy duty and the transparency rules, which did not move, probably are.</p></div>
<h2>The penalties are unchanged, and they are large</h2>
<p>None of this softened the enforcement side. Article 99 still sets fines of up to 35 million euro or 7 percent of global annual turnover for prohibited practices, up to 15 million euro or 3 percent for most other breaches, and up to 7.5 million euro or 1 percent for supplying wrong information to authorities. You can read the exact tiers on the <a href="https://artificialintelligenceact.eu/article/99/">AI Act Explorer</a>. For a mid-sized firm, the percentage figures are the ones that bite.</p>
<h2>What a deployer should actually do</h2>
<p>Strip away the calendar noise and the to-do list is short. Work out where AI is already being used across the organisation, including the tools nobody officially bought. Sort each use into a bucket: banned, high-risk, limited-risk (which mostly means transparency duties), or minimal. Get the AI literacy training done, because it is already overdue and it is one of the cheapest boxes to tick. And for anything that even might be high-risk, name an owner now, because December 2027 arrives faster than it sounds when standards, documentation and human-oversight processes all have to be in place first.</p>
<p>One quieter point. A lot of the deployer burden is about evidence: keeping logs, pinning the exact model version you assessed, being able to say what the system does with an input. All of that is easier when the model runs on infrastructure you control rather than behind a vendor's API that can change under you. Sovereignty and compliance are not the same thing, but they rhyme.</p>
<p>This is general information, not legal advice. The omnibus text was still being finalised in the Official Journal as this went out, so confirm the final regulation number and any national guidance with your counsel before you act on it.</p>]]></content:encoded>
    </item>
    <item>
      <title>Why we put a datacenter in a locker</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/why-we-put-a-datacenter-in-a-locker/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/why-we-put-a-datacenter-in-a-locker/</guid>
      <pubDate>Thu, 02 Jul 2026 09:00:00 GMT</pubDate>
      <category>Blog</category>
      <category>Product</category>
      <description>The thinking behind Kaldryn One, why the future of business AI for most companies is a quiet black box in the corner of the office, not a bigger cloud bill.</description>
      <content:encoded><![CDATA[<p>When we demo Kaldryn One, the most common reaction isn't a question about tokens per second or context windows. It's someone pointing at the box and asking: <em>“That's it?”</em></p>
<p>That's it. A locker-sized black appliance, quieter than the office fridge, plugged into a normal wall outlet and your network switch. Inside it: your company's entire AI capability. Chat over your own documents, more than two hundred curated open models, an OpenAI-compatible API, and not a single byte leaving the building.</p>
<h2>The problem with “just use the cloud”</h2>
<p>For a lot of teams, cloud AI is genuinely fine. But every conversation we had with hospitals, law firms, banks and public agencies across the Nordics followed the same arc: people were getting real value out of consumer AI tools, leadership wanted that capability properly, and every path to it ran through somebody else's infrastructure, under somebody else's terms.</p>
<ul>
<li>The pricing scales with headcount. Per-seat AI subscriptions are a tax on growing.</li>
<li>The data leaves. However good the contract is, a prompt typed into a cloud assistant physically departs your premises.</li>
<li>The dependency compounds. Models get deprecated, terms get revised, prices change, and your workflows are built on all of it.</li>
</ul>
<p>For regulated organisations, the second point alone often ends the discussion. Our answer is not “trust us instead”. It's: don't send the data anywhere at all.</p>
<h2>Constraints make products</h2>
<p>We gave ourselves three constraints for the appliance, and they shaped everything:</p>
<ol>
<li><strong>It has to live in an office, not a server room.</strong> That set the acoustic budget (under 38 dB, quieter than normal conversation), the power budget (a regular wall outlet, no three-phase), and the size (locker, not rack).</li>
<li><strong>An IT team of one has to run it.</strong> Two cables, power and ethernet. It arrives pre-configured, indexes your documents on-box, and serves your team the same day.</li>
<li><strong>It has to keep working with the internet unplugged.</strong> Not as a party trick, as the operating assumption. Updates are something you choose to apply, not something that happens to you.</li>
</ol>
<div data-callout><p>The test we kept coming back to: if you pull the network cable out of the wall, what still works? For Kaldryn One the answer is: everything. That's the difference between owning your AI and renting it.</p></div>
<h2>What “a datacenter” actually means for a 30-person firm</h2>
<p>The phrase sounds grandiose, but look at what a small organisation actually needs from AI infrastructure: an assistant that knows its documents, a private API for the tools it builds, and governance a lawyer can sign off. That's not a hyperscale problem. It fits in a box, GPUs, storage, models, retrieval, admin console, audit log, all of it.</p>
<p>One box for a thirty-person firm. Racks of them for a data center. The platform is the same either way, which is exactly the point. Sovereignty shouldn't be an enterprise luxury.</p>
<h2>What's next</h2>
<p>We're publishing more of the thinking behind the platform here on the blog, the engineering trade-offs, what runs well on one box in 2026, and what we're hearing from the regulated side of the market. For the deeper technical material, the <a href="/research/">research hub</a> has the reference architecture and compliance white papers. And if you'd rather just see the box: <a href="/kaldryn-one/">meet Kaldryn One</a>.</p>]]></content:encoded>
    </item>
    <item>
      <title>The EU AI Act timeline: what applies now and what's coming</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/eu-ai-act-timeline-for-deployers/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/eu-ai-act-timeline-for-deployers/</guid>
      <pubDate>Wed, 01 Jul 2026 09:00:00 GMT</pubDate>
      <category>EU regulations</category>
      <category>EU AI Act</category>
      <description>A plain-language walkthrough of the AI Act's phased application dates, what organisations deploying AI must already comply with, and what lands next.</description>
      <content:encoded><![CDATA[<p>The AI Act is often discussed as if it were one deadline. It isn't, it's a staircase. The regulation entered into force on 1 August 2024, and its obligations switch on in phases over several years. If your organisation <em>uses</em> AI (in AI Act terms, if you're a <strong>deployer</strong>), here is the staircase in plain language.</p>
<h2>The phased timeline</h2>
<table><thead><tr><th>Date</th><th>What starts applying</th><th>Who feels it</th></tr></thead><tbody>
<tr><td>1 Aug 2024</td><td>Regulation enters into force</td><td>Everyone (clock starts)</td></tr>
<tr><td>2 Feb 2025</td><td>Prohibited practices banned; AI-literacy duty (Art. 4)</td><td>All providers &amp; deployers</td></tr>
<tr><td>2 Aug 2025</td><td>General-purpose AI (GPAI) model obligations; governance &amp; penalties framework</td><td>Model providers, primarily</td></tr>
<tr><td>2 Aug 2026</td><td>The bulk of the Act, including high-risk systems (Annex III) and the Art. 50 transparency duties</td><td>Providers &amp; deployers broadly</td></tr>
<tr><td>2 Aug 2027</td><td>High-risk rules for AI embedded in regulated products (Art. 6(1))</td><td>Product manufacturers</td></tr>
</tbody></table>
<h2>What you should already be doing</h2>
<p>Two obligations have applied to essentially every organisation using AI since February 2025:</p>
<ul>
<li><strong>Avoiding prohibited practices.</strong> The banned list (manipulative techniques, social scoring, most emotion recognition in workplaces, and similar) applies regardless of company size or sector.</li>
<li><strong>AI literacy (Article 4).</strong> Staff who operate or use AI systems must have a sufficient level of AI literacy. This is a real, current obligation, and one of the easiest to evidence with a training module and a record of completion.</li>
</ul>
<h2>What lands on 2 August 2026</h2>
<p>The big one. From this date the high-risk regime applies to Annex III use cases, think employment and HR screening, credit scoring, essential services, education, and most public-sector decision support. Deployers of high-risk systems get concrete duties: use the system per its instructions, ensure human oversight, keep logs, feed incidents back to the provider, and in some cases run a fundamental-rights impact assessment (FRIA).</p>
<p>Alongside it, the Article 50 transparency duties apply: people must be told when they're interacting with an AI system, and AI-generated content must be disclosed in a machine-readable way where required.</p>
<div data-callout><p>A common misconception: “we only use an internal chatbot, so the AI Act isn't about us.” Mostly true for the high-risk regime, an internal drafting assistant is normally not Annex III, but the AI-literacy duty, the prohibited-practices ban and (from Aug 2026) the transparency rules still apply to you.</p></div>
<h2>Where self-hosting changes the picture</h2>
<p>The AI Act doesn't care where your servers are, obligations follow the use case, not the hosting model. What self-hosting <em>does</em> change is how much of the evidence chain you control. When the model runs on your hardware, you can pin the exact model version you assessed, produce your own logs without waiting on a vendor, and answer "what does the system actually do with inputs" from first-hand knowledge rather than a sub-processor list. That's why Kaldryn ships the deployer tooling, transparency banner, AI-literacy module, FRIA/DPIA templates and an immutable audit log, in the box.</p>
<h2>Sensible next steps</h2>
<ol>
<li>Inventory your AI use and sort it: prohibited (stop), high-risk (prepare for Aug 2026), limited-risk (transparency), minimal (document and move on).</li>
<li>Stand up AI-literacy training now if you haven't, it's already due.</li>
<li>For anything potentially high-risk, assign an owner for the deployer duties and start the FRIA groundwork this year.</li>
</ol>]]></content:encoded>
    </item>
    <item>
      <title>Why European companies stopped trusting the cloud with everything</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/why-europe-stopped-trusting-the-cloud-with-everything/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/why-europe-stopped-trusting-the-cloud-with-everything/</guid>
      <pubDate>Sat, 27 Jun 2026 09:00:00 GMT</pubDate>
      <category>Blog</category>
      <category>Field notes</category>
      <description>This did not come from nowhere. A court ruling, a foreign law, five years of unease and, in 2026, governments actually moving. A short history of the shift.</description>
      <content:encoded><![CDATA[<p>If you had said, a few years ago, that European governments would start ripping American software out of their offices on principle, you would have sounded paranoid. In 2026 it is just the news. This shift did not appear from nowhere, and it is worth understanding the arc, because it explains a lot about why so many European companies now draw a line around their most sensitive data.</p>
<h2>It started in a courtroom</h2>
<p>The turning point most people can name is <a href="https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:62018CJ0311">Schrems II</a>, the 2020 ruling from the EU's top court that struck down the main legal bridge for sending personal data to the United States. The court's problem was structural, not petty: US surveillance law let American agencies reach data held by American providers, and a European whose data got caught up had no real way to object. That mismatch has never been fully resolved. Each replacement arrangement has been challenged, and the current one sits on the same fault line as the two that fell before it.</p>
<h2>Then people read the fine print</h2>
<p>Underneath the court case was a quieter realisation about the US CLOUD Act. In plain terms, it can compel American providers to hand over data they hold, even when that data is stored on servers physically located in Europe. Read that twice. Your data can sit in a data centre in Frankfurt or Stockholm and still be reachable under a foreign law, because the company holding it answers to a foreign government. For a hospital, a ministry or a bank, that is not a comfortable footnote. It is the whole problem.</p>
<h2>2026 is when talk turned into moving vans</h2>
<p>For years this was a debate. Recently it became action. In June 2025 Denmark's digitalisation ministry announced it was moving off Microsoft software toward open-source alternatives, citing sovereignty and cost, as <a href="https://therecord.media/denmark-digital-agency-microsoft-digital-independence">reported at the time</a>. The German state of Schleswig-Holstein set out to shift tens of thousands of machines to Linux and open tools, as <a href="https://www.theregister.com/2025/10/15/schleswig_holstein_open_source/">the trade press covered</a>. And in May 2026 the Swedish government adopted its <a href="https://www.regeringen.se/pressmeddelanden/2026/05/ny-molnpolicy-ska-bidra-till-okad-digital-suveranitet-i-offentlig-forvaltning/">first national cloud policy</a>, stating outright that the country is too dependent on suppliers outside the EU. The European Commission followed with a proposed Cloud and AI Development Act built around the same worry.</p>
<div data-callout><p>None of this is anti-American, whatever the headlines suggest. It is about control. A country or a company that cannot function without a supplier it does not govern has handed away something important, and after five years of unease, a lot of Europe decided to take some of it back.</p></div>
<h2>Where AI walked straight into it</h2>
<p>Here is the irony. Just as Europe grew wary of sending sensitive data to foreign clouds, along came a technology that wants to send it all: prompts, documents, context, the lot, straight to a provider's model. Adopting cloud AI carelessly undoes years of careful work on data transfers in a single procurement.</p>
<p>Which is why running AI locally is not really a fringe stance any more. It is the same instinct that is moving government offices to open source, expressed at the level of a single tool. Keep the capability, keep the data at home. We happen to be a Swedish company that built AI this way from the start, so we are not surprised the continent is arriving here. We just got here early.</p>]]></content:encoded>
    </item>
    <item>
      <title>The Cyber Resilience Act: security duties for anything with software</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/cyber-resilience-act-software-security-duties/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/cyber-resilience-act-software-security-duties/</guid>
      <pubDate>Wed, 24 Jun 2026 09:00:00 GMT</pubDate>
      <category>EU regulations</category>
      <category>Sector rules</category>
      <description>The Cyber Resilience Act sets baseline security rules for products with digital elements. What it requires, when it bites, and how it touches AI systems.</description>
      <content:encoded><![CDATA[<p>The <a href="https://eur-lex.europa.eu/eli/reg/2024/2847/oj">Cyber Resilience Act</a>, Regulation (EU) 2024/2847, is the EU's answer to a simple, embarrassing fact: for years you could sell software and connected devices in Europe with no baseline security obligations at all. That changes. The CRA entered into force on 10 December 2024, and its main requirements apply from 11 December 2027, with the vulnerability-reporting duties arriving earlier, on 11 September 2026.</p>
<h2>What it covers, and what it asks</h2>
<p>The CRA applies to products with digital elements, which is a deliberately wide net: hardware and software that can connect to a device or a network, from smart home gadgets to operating systems, firewalls and password managers. If you place such a product on the EU market, you inherit a set of duties. The <a href="https://digital-strategy.ec.europa.eu/en/policies/cyber-resilience-act">Commission's CRA pages</a> are the reference.</p>
<p>The core requirements are the sort of thing security people have asked for politely for a decade and are now getting in law. Products have to be secure by design and secure by default. They must ship without known exploitable vulnerabilities. Makers have to run a real vulnerability-handling process, with coordinated disclosure and free security updates across a defined support period. And they have to produce a software bill of materials, an SBOM, so buyers can actually see what is inside the thing they bought.</p>
<div data-callout><p>The SBOM requirement is quietly a big deal for AI. If you are shipping or running a system built from open-source components, model runtimes and libraries, an SBOM forces you to know your own ingredients. That is exactly the kind of supply-chain hygiene that gets skipped when everything is a black-box API you rent by the token.</p></div>
<h2>How it touches AI systems</h2>
<p>An AI product is still a product. If it has digital elements and it goes on the market, the CRA's security-by-design and vulnerability-handling duties apply to it like anything else, sitting alongside whatever the AI Act asks. There are carve-outs, notably for non-commercial open-source software and for products already covered by sectoral rules like medical devices, but the general software case is squarely in scope.</p>
<p>There is a reporting edge worth flagging for anyone shipping software. From 11 September 2026, makers have to report actively exploited vulnerabilities and severe incidents to ENISA and national response teams on a tight clock, with an early warning inside 24 hours. That is a real operational obligation, not a paperwork one.</p>
<h2>The self-hosting angle</h2>
<p>The CRA does not care whether your AI runs in the cloud or on your own metal. Its duties attach to products, not deployment locations. But the philosophy behind it, know what your software is made of, patch it, do not ship known holes, is far easier to honour when you control the stack. When you run an AI system in-house, you can produce a real SBOM, apply updates on your schedule, and account for every component. When your AI capability is a remote API, your supply-chain visibility ends at the vendor's edge, and you are trusting their CRA compliance as much as your own.</p>
<p>Regulators keep circling the same idea from different directions. The <a href="https://www.enisa.europa.eu/">EU cybersecurity agency ENISA</a> has documented for years how supply-chain attacks reach customers through their suppliers rather than head-on. The CRA, DORA and NIS2 all push the same way: know your dependencies, shorten your chains, own your risk. Keeping AI on-premise is one concrete way to do all three at once.</p>
<p>This is general information and not legal advice. The CRA's scope, classes and timelines have real nuance, so confirm how it applies to your specific product with a qualified adviser.</p>]]></content:encoded>
    </item>
    <item>
      <title>The questions every IT lead asks before going on-prem with AI</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/questions-it-leads-ask-before-onprem-ai/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/questions-it-leads-ask-before-onprem-ai/</guid>
      <pubDate>Wed, 24 Jun 2026 09:00:00 GMT</pubDate>
      <category>Blog</category>
      <category>AI in practice</category>
      <description>Hardware, models, maintenance, updates and what happens when the internet goes down, the honest answers to the questions we hear in every Kaldryn evaluation.</description>
      <content:encoded><![CDATA[<p>Every serious evaluation of on-premise AI lands on the same handful of questions. They're good questions, the kind an IT lead <em>should</em> ask before putting a new box on the network. Here are the ones we hear most, answered the way we answer them in the room.</p>
<h2>“Do we need a GPU cluster and a DevOps team?”</h2>
<p>No. This is the assumption that stops most teams from even evaluating on-prem, and it's a few years out of date. A modern open-weights model, quantized and served well, delivers a genuinely useful assistant on a single appliance-class machine. Kaldryn installs on a fresh Ubuntu server in under ten minutes, no Kubernetes, no platform team. If you can rack a NAS, you can run your own AI.</p>
<h2>“Are open models actually good enough?”</h2>
<p>For the everyday work an internal assistant does, drafting, summarising, answering questions over your own documents, yes, and the gap to frontier cloud models has narrowed every quarter. The trick is matching the model to the job: a fast small model for chat and retrieval, a bigger one for heavy drafting. Kaldryn ships 200+ curated models with signed weights, so choosing isn't a research project.</p>
<div data-callout><p>The question that matters isn't “is this the best model in the world?”, it's “is this model, running over <em>our</em> documents, more useful than a frontier model that isn't allowed to see them?” For regulated teams, the second one wins.</p></div>
<h2>“Who maintains it?”</h2>
<p>The same person who maintains the rest of your infrastructure, which is to say, not much of anyone, most weeks. The platform monitors itself, the admin console shows what your users are doing (and what they're trying to do, the DLP log is illuminating), and updates are explicit: you decide when, they're signed, and you can defer them indefinitely on air-gapped deployments.</p>
<h2>“What happens when the internet goes down?”</h2>
<p>Nothing. That's the whole design. Retrieval, chat, the API, everything runs on-box. An internet outage that takes the office offline takes your cloud AI subscriptions with it; the appliance in the corner doesn't notice.</p>
<h2>“What does it cost, really?”</h2>
<p>Owned hardware is a fixed cost; per-seat cloud AI is a linear one. The crossover point depends on your headcount and usage, but the shape of the curve is the part to internalise: an appliance costs the same whether ten people use it or two hundred. Past a modest team size, “unlimited users, just electricity” stops sounding like marketing and starts being the spreadsheet.</p>
<h2>“How do we prove to compliance that nothing leaves?”</h2>
<p>This is where on-prem flips from being the cautious choice to the easy one. The answer isn't a vendor questionnaire and a sub-processor list, it's a firewall rule and a packet capture. Kaldryn seals its network perimeter on boot and ships the evidence tooling (immutable audit log, SIEM export, isolation checks) to demonstrate it. Your DPO gets an architecture where the sensitive question, <em>where does the data go?</em>, has a one-word answer: nowhere.</p>
<p>Evaluating a deployment for your own team? The <a href="/research/on-premise-ai-reference-architecture/">reference architecture</a> covers the technical detail, and <a href="/eu-regulations/">our EU regulation briefings</a> cover the compliance side.</p>]]></content:encoded>
    </item>
    <item>
      <title>Sovereign AI: What It Means and Why Regulated Enterprises Need It</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/research/what-is-sovereign-ai/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/research/what-is-sovereign-ai/</guid>
      <pubDate>Wed, 24 Jun 2026 09:00:00 GMT</pubDate>
      <category>Research &amp; white papers</category>
      <category>Sovereign AI</category>
      <description>A plain-language primer on sovereign AI, what it is, how it differs from cloud AI, and why regulated organisations are bringing their models in-house.</description>
      <content:encoded><![CDATA[<p>“Sovereign AI” has gone from a policy buzzword to a procurement requirement in under two years. The short definition: an organisation has sovereign AI when it controls the three things that matter, the model, the data it processes, and the infrastructure it runs on, without depending on a third party who can change the terms, read the inputs, or switch the service off.</p>
<p>That sounds obvious until you look at how most enterprises consume AI today. A prompt typed into a cloud assistant leaves the building, is processed on infrastructure the customer never sees, under a contract that can be revised, and is often retained for some period for abuse-monitoring or training. For a marketing team that trade-off is fine. For a hospital, a law firm, a bank, or a defence contractor, it is frequently a non-starter.</p>
<h2>The three layers of sovereignty</h2>
<p>It helps to be precise, because vendors use the word loosely. True sovereignty has three independent layers, and a deployment is only as sovereign as its weakest one:</p>
<ul><li>Data sovereignty, your inputs, outputs and documents physically stay on hardware you control and are never transmitted to a third party.</li><li>Model sovereignty, you can inspect, pin, and keep running the exact model you deployed, regardless of a vendor deprecating or altering it.</li><li>Operational sovereignty, the system keeps working without phoning home, even fully air-gapped, and you decide when and whether to update it.</li></ul>
<p>A “private” cloud endpoint typically gives you the first layer on paper and none of the others. A downloaded open-weights model gives you model sovereignty but leaves you to solve operations, governance and security yourself. Sovereign AI as a product category is the attempt to deliver all three at once.</p>
<h2>Why now</h2>
<p>Three forces are converging. First, regulation: the EU AI Act, GDPR enforcement, sector rules like DORA and NIS2, and national data-residency expectations have made “where does the data go?” a board-level question. Second, capability: open-weights models are now good enough that a self-hosted assistant is genuinely useful for drafting, summarising and document Q&amp;A, the gap to frontier cloud models has narrowed for everyday enterprise tasks. Third, economics: per-seat cloud AI pricing scales linearly with headcount, while owned hardware is a fixed cost that the organisation keeps.</p>
<blockquote><p>“The question has quietly flipped. It used to be “why would we run AI ourselves?” Now, for regulated buyers, it is “why are we still sending this data to someone else?””</p></blockquote>
<h2>What sovereignty is not</h2>
<p>Sovereignty is not isolationism, and it does not mean giving up a modern experience. A well-built sovereign platform still offers streaming chat, retrieval over your own documents, image understanding, an admin console and an API, it simply runs them on metal you control. Nor does sovereignty mean “air-gapped or nothing”: most organisations run connected for convenience and reserve full network isolation for their most sensitive workloads. The point is that the choice is yours to make and to change.</p>
<div data-callout><p>A useful test: if your AI vendor went out of business tomorrow, or doubled their price, or was compelled to hand over logs, would your AI capability survive intact? If the answer is no, you do not yet have sovereign AI.</p></div>
<h2>Who needs it first</h2>
<p>Adoption is led by the organisations with the least room to be wrong about data handling:</p>
<ul><li>Healthcare, patient data under GDPR Article 9, MDR/IVDR and the emerging European Health Data Space.</li><li>Legal, privilege and client confidentiality that cannot survive a third-party processor.</li><li>Finance, DORA, supervisory expectations and the sensitivity of transaction and client data.</li><li>Defence and public sector, classification, national sovereignty and air-gapped environments by default.</li></ul>
<p>What these sectors share is that the cost of a data-handling mistake is measured in licences, lawsuits and lives, not churn. For them sovereignty is the baseline, and a capable assistant on top of it is the prize.</p>
<h2>Getting there</h2>
<p>The practical path is less daunting than it sounds. Sovereign AI runs on commodity GPU servers or even capable CPU nodes for smaller models; it installs in minutes rather than requiring a Kubernetes platform team; and the governance, isolation and compliance tooling that regulated buyers need can ship in the box rather than being assembled from a dozen projects. That is the bet Kaldryn is built on, and the rest of this research hub goes deeper on each layer, from reference architectures to EU AI Act compliance.</p>]]></content:encoded>
    </item>
    <item>
      <title>On-Premise AI: A Reference Architecture for Air-Gapped LLM Deployment</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/research/on-premise-ai-reference-architecture/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/research/on-premise-ai-reference-architecture/</guid>
      <pubDate>Sat, 20 Jun 2026 09:00:00 GMT</pubDate>
      <category>Research &amp; white papers</category>
      <category>Architecture</category>
      <description>A practical reference architecture for deploying large language models on-premise and air-gapped, network topology, GPU sizing, model serving and zero-egress design.</description>
      <content:encoded><![CDATA[<p>This white paper sketches a reference architecture for running a production large-language-model assistant entirely inside your own perimeter, including the fully air-gapped case. It is deliberately vendor-neutral in its principles, with notes on how Kaldryn implements each piece.</p>
<h2>Design goals</h2>
<ul><li>Zero data egress, no inference input, output or document ever leaves the perimeter.</li><li>No phone-home, licensing, updates and operation work without outbound internet.</li><li>Single-server start, cluster-ready growth, scale out without rebuilding.</li><li>Defence in depth, isolation enforced at multiple independent layers.</li><li>Operable by a small team, no dedicated platform or DevOps org required.</li></ul>
<h2>The five planes</h2>
<p>It is useful to decompose the appliance into five planes, each with a clear responsibility:</p>
<ol><li>Serving plane, the model runtime (e.g. vLLM and Ollama behind a load balancer) with tier-aware routing across CPU and GPU nodes.</li><li>Retrieval plane, document ingestion, embeddings, a vector store and security-trimmed RAG so answers cite your own files.</li><li>Identity plane, SSO/SAML, group-to-role mapping and just-in-time provisioning, with metadata upload for air-gapped IdPs.</li><li>Governance plane, DLP scanning, an immutable audit log, and policy enforcement on every inbound and outbound message.</li><li>Observability plane, latency percentiles, tokens/sec, queue depth and KV-cache utilisation per tier.</li></ol>
<div data-callout><p>Keeping these planes distinct matters operationally: you can scale the serving plane (add GPUs) independently of the retrieval plane (add storage and embedding throughput) as usage grows.</p></div>
<h2>Network topology</h2>
<p>For the air-gapped deployment, the topology is intentionally boring. Clients reach the appliance over the internal network only; the appliance has no default route to the internet. Where an organisation wants connected convenience for some workloads and isolation for others, the recommended pattern is a single appliance running in data-isolation mode with egress sealed, rather than two separate systems to keep in sync.</p>
<p>Air-gapping should never rely on one toggle. We recommend three independent layers:</p>
<ul><li>Application layer, outbound networking disabled at the code level so the software cannot initiate external calls.</li><li>Host layer, an egress firewall (e.g. nftables) that drops outbound traffic regardless of application behaviour.</li><li>Network layer, no route to the internet from the appliance's VLAN, verified by the customer's own controls.</li></ul>
<h2>GPU sizing</h2>
<p>Sizing is driven by three variables: the parameter count and quantisation of the models you intend to serve, your target concurrency, and your latency budget (time-to-first-token and tokens/sec). As rough guidance, always validate with your own benchmark on your own hardware:</p>
<table><thead><tr><th>Deployment</th><th>Typical hardware</th><th>Good for</th></tr></thead><tbody><tr><td>Pilot / small team</td><td>1 server, 1-2 modern data-centre GPUs</td><td>Up to a few dozen active users, mid-size models</td></tr><tr><td>Department</td><td>1-2 servers, 4-8 GPUs</td><td>Hundreds of users, larger models, RAG at scale</td></tr><tr><td>Organisation</td><td>Multi-node GPU cluster, load-balanced</td><td>500+ users, tiered model routing, high concurrency</td></tr></tbody></table>
<p>The important architectural property is that moving between these tiers is a capacity change, not a re-platforming. Tier-aware routing lets you add nodes and point heavier workloads at GPU tiers while lighter ones stay on cheaper hardware.</p>
<h2>Model serving and provenance</h2>
<p>Each model should be verified by content hash on first load and pinned thereafter, so the weights you reviewed are the weights that run. A provenance grade, covering license, source, stability and integrity, lets administrators make informed choices from a curated catalogue rather than pulling arbitrary weights from the internet. We cover this in depth in our note on model provenance.</p>
<h2>Governance on the hot path</h2>
<p>DLP and audit are not bolt-ons; they sit on the message hot path. A regex-and-rules scan runs on every inbound and outbound message to block, redact or audit PII, secrets and network data, and every action is written to an immutable, SIEM-exportable log. Because nothing leaves the perimeter, this audit trail is also the organisation's complete record for eDiscovery and supervisory review.</p>
<h2>Deployment and day-two operations</h2>
<p>The reference deployment is a single signed, checksummed installer that carries every image, model and dependency, runs preflight checks, stands up services and warms a default model, no manual wiring. For air-gapped sites, updates arrive as signed media rather than over the network, and restores run detached so the appliance cannot be bricked by a bad upgrade. The goal throughout is that a small IT team, not a platform org, can own it.</p>
<p>Want the full version of this reference architecture, including diagrams and a sizing worksheet? Talk to our team, we provide it under NDA for evaluation.</p>]]></content:encoded>
    </item>
    <item>
      <title>Digital sovereignty in the EU: a 2026 field guide</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/research/digital-sovereignty-eu-2026-field-guide/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/research/digital-sovereignty-eu-2026-field-guide/</guid>
      <pubDate>Thu, 18 Jun 2026 09:00:00 GMT</pubDate>
      <category>Research &amp; white papers</category>
      <category>Sovereign AI</category>
      <description>Europe spent 2025 and 2026 waking up to its dependence on a few foreign cloud providers. A field guide to the laws, the money and the moves that follow.</description>
      <content:encoded><![CDATA[<p>For most of the past decade, "digital sovereignty" was a phrase for conference panels. In 2025 and 2026 it turned into policy, budgets and actual migrations. If you sell, buy or run technology in Europe, it is worth understanding what changed, because the ground moved under a lot of assumptions about where data and computing are allowed to live.</p>
<h2>The number that started the conversation</h2>
<p>Estimates vary, but the shape is consistent and uncomfortable. Something like ninety percent of Europe's digital infrastructure runs on non-European, mostly American, providers, and the three big US hyperscalers hold roughly two-thirds of the European cloud market. The European providers' share actually fell over recent years. Analyses like the <a href="https://www.bertelsmann-stiftung.de/en/our-projects/reframetech-algorithmen-fuers-gemeinwohl/project-news/eurostack-a-european-alternative-for-digital-sovereignty">EuroStack work from the Bertelsmann Stiftung</a> lay out the dependence in detail. Once you see the figures, a lot of European policy over the past two years reads as a single reaction to them.</p>
<h2>The laws catching up</h2>
<p>Several legal threads pull in the same direction. The <a href="https://eur-lex.europa.eu/eli/reg/2023/2854/oj/eng">EU Data Act</a>, applicable since September 2025, is dismantling the fees and technical obstacles that lock customers into a cloud provider, and it obliges providers to guard non-personal data against unlawful foreign-government access. The GDPR's transfer rules, shaped by the Schrems II ruling, keep the risk of sending personal data abroad firmly on the table. And in June 2026 the Commission proposed a <a href="https://digital-strategy.ec.europa.eu/en/policies/cloud-and-ai-development-act">Cloud and AI Development Act</a>, part of a wider tech-sovereignty package, whose own framing is blunt: over-reliance on non-EU cloud providers is a significant risk to Europe's digital autonomy. The proposal talks about tripling EU data-centre capacity and defining tiers of cloud sovereignty, the strongest of which rules out third-country corporate control for the most sensitive public workloads.</p>
<h2>Sweden puts it in writing</h2>
<p>This is not only a Brussels story. On 28 May 2026 the Swedish government adopted its <a href="https://www.regeringen.se/pressmeddelanden/2026/05/ny-molnpolicy-ska-bidra-till-okad-digital-suveranitet-i-offentlig-forvaltning/">first national cloud policy for public administration</a>, stating in plain Swedish that the country is heavily dependent on suppliers outside the EU. The policy does not ban American cloud. It requires risk assessments, open standards, clear contracts and exit plans, so that a public body can actually leave a provider if it needs to. The direction is unmistakable even where the language is careful.</p>
<div data-callout><p>Notice the recurring word in all of this: exit. The Data Act removes exit fees. The Swedish policy demands exit plans. The EU sovereignty tiers are about not being trapped. Regulators have concluded that the ability to leave is the thing that was missing. The cleanest version of an exit plan is the arrangement you never have to exit, because the data was in your building the whole time.</p></div>
<h2>From policy to actual migrations</h2>
<p>The most telling signal is that organisations stopped waiting for the laws to finish. In June 2025 Denmark's digitalisation ministry announced a move from Microsoft software toward open-source alternatives, on explicit sovereignty grounds, as <a href="https://therecord.media/denmark-digital-agency-microsoft-digital-independence">reported at the time</a>, with Copenhagen citing a Microsoft licence bill that had climbed sharply over five years. In Germany, the state of Schleswig-Holstein set out to move tens of thousands of machines off Windows and Office to Linux and open-source tools, projecting real annual savings, as <a href="https://www.theregister.com/2025/10/15/schleswig_holstein_open_source/">covered in the trade press</a>. These are not startups making a point. They are governments moving payroll and email.</p>
<h2>Where AI fits</h2>
<p>AI sharpens every part of this. An assistant that sends prompts and documents to a foreign-controlled cloud is the exact dependence the sovereignty debate is about, applied to your most sensitive internal knowledge. Running the model on your own hardware is not a clever hedge against the trend. It is the trend, expressed at the level of a single deployment. When the computation happens where the data already lives, there is no provider to be locked into, no border to defend, and no foreign statute that reaches the server.</p>
<p>None of this means American cloud is going away, and this guide is not a prediction that it should. It is a map of a real shift in what European institutions consider prudent. For a Swedish company that has always built AI to run inside the customer's walls, the map mostly confirms a bet we made early. The rest of the continent is now describing, in law and in budgets, the thing we build.</p>]]></content:encoded>
    </item>
    <item>
      <title>GDPR and self-hosted LLMs: the questions your DPO will ask</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/gdpr-self-hosted-llms-dpo-questions/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/gdpr-self-hosted-llms-dpo-questions/</guid>
      <pubDate>Thu, 18 Jun 2026 09:00:00 GMT</pubDate>
      <category>EU regulations</category>
      <category>GDPR</category>
      <description>Lawful basis, minimisation, the right to erasure and embeddings, DPIAs, how the GDPR conversation changes when the model runs on your own hardware.</description>
      <content:encoded><![CDATA[<p>Put an LLM in front of employees and, sooner or later, personal data goes through it, in prompts, in retrieved documents, in generated drafts. That makes the deployment a GDPR matter, whoever hosts it. But the conversation with your DPO goes very differently depending on where the model runs. Here are the questions that come up, and how self-hosting changes the answers.</p>
<h2>“Where does the data go, and who is processing it?”</h2>
<p>With cloud AI this question spawns a diagram: your organisation (controller), the AI vendor (processor), their infrastructure provider (sub-processor), possibly a transfer mechanism if anything touches a third country. Each hop needs contracts, assessment and monitoring.</p>
<p>With a self-hosted deployment the diagram collapses: the data goes to a machine you own, processed by software you operate. There is no processor for the inference step, no sub-processor chain, no international transfer. Whole categories of GDPR homework, DPAs for the AI vendor, transfer impact assessments, sub-processor change notifications, simply don't arise for the core processing.</p>
<h2>“What's our lawful basis, and did we minimise?”</h2>
<p>Hosting doesn't change this one: you still need a lawful basis (typically legitimate interests for internal productivity tooling, with a documented balancing test) and you still need data-minimisation discipline. What helps in practice is <em>control at the gateway</em>: DLP rules that stop obvious over-sharing (personnummer, card numbers, patient identifiers) before it reaches the model, and retention settings you define rather than inherit from a vendor's policy.</p>
<h2>“Can we actually delete? What about the embeddings?”</h2>
<p>The Article 17 question is where LLM deployments get genuinely technical. A person's data can live in five places: chat histories, uploaded documents, the vector index built from those documents (embeddings), any fine-tuned model trained on them, and backups. An honest erasure story has to reach all five.</p>
<div data-callout><p>Ask any AI vendor, including us, the same question: “When I delete a document, what happens to the vectors derived from it, and how long until backups age out?” If the answer isn't specific, the erasure story isn't real. On Kaldryn, erasure propagates across conversations, documents, embeddings, fine-tunes and backups, and the audit log records the propagation.</p></div>
<h2>“Do we need a DPIA?”</h2>
<p>Usually yes, and you should want one anyway, a DPIA is where the design decisions (what data the assistant can see, who can query what, retention, access) get made deliberately instead of by default. The EDPB's work on AI models (notably its Opinion 28/2024 on personal data in model training) is worth your DPO's time here; among other things it underlines that whether a model itself contains personal data is a case-by-case question. Self-hosting makes the mitigations section of the DPIA notably shorter: no third-party access, no transfers, physical control of the hardware.</p>
<h2>“What do we tell employees?”</h2>
<p>Transparency duties apply regardless of hosting: update your internal privacy notice, be clear about what's logged (admins can typically see usage), and set rules for what belongs in prompts. An internal AI-use policy plus the AI-literacy training you already owe under the AI Act covers most of it.</p>
<p>The short version for the steering meeting: GDPR applies either way, but self-hosting converts your hardest external-dependency questions into internal engineering questions, and internal engineering questions have answers you can verify.</p>]]></content:encoded>
    </item>
    <item>
      <title>The EU AI Act and Self-Hosted LLMs: A Compliance Checklist</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/research/eu-ai-act-self-hosted-llms-checklist/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/research/eu-ai-act-self-hosted-llms-checklist/</guid>
      <pubDate>Tue, 16 Jun 2026 09:00:00 GMT</pubDate>
      <category>Research &amp; white papers</category>
      <category>Compliance</category>
      <description>How the EU AI Act maps onto self-hosted and on-premise LLM deployments, a practical, plain-language checklist for legal, compliance and security teams.</description>
      <content:encoded><![CDATA[<p>This is a practical orientation for legal, compliance and security teams weighing a self-hosted AI deployment against the EU AI Act and GDPR. It is not legal advice, your EU-qualified counsel does the final review, but it maps the obligations that most often surprise teams onto concrete, on-premise controls.</p>
<div data-callout><p>Kaldryn ships the tooling and the documentation scaffolding for these obligations. It does not replace a lawyer. Treat this checklist as a starting map, not a guarantee.</p></div>
<h2>First, self-hosting is not an exemption</h2>
<p>A common misconception is that running AI on your own hardware puts you outside the EU AI Act. It does not. The Act regulates the deployment and use of AI systems based on risk, not on where the GPUs sit. What self-hosting changes is the difficulty of compliance: when data never leaves your perimeter, a whole class of obligations around transfer, processing and residency become architectural facts rather than contractual promises you have to extract from a vendor.</p>
<h2>Know your risk tier</h2>
<p>The Act is risk-tiered. Most enterprise assistant use, drafting, summarising, internal document Q&amp;A, falls outside the high-risk categories, but you must confirm that for your specific use cases. Deploying AI in areas such as recruitment, credit, biometric identification or critical infrastructure can pull you into the high-risk regime with substantially more obligations. Step one is always a documented classification of each use case.</p>
<h2>The transparency obligation (Article 50)</h2>
<p>Users must know when they are interacting with an AI system, and AI-generated content must be marked as such where required. In practice this means a transparency notice in the assistant UI, locale-aware and customisable, plus clear labelling of AI output. This is straightforward to satisfy but easy to forget.</p>
<h2>AI literacy (Article 4)</h2>
<p>From 2025 the Act expects organisations to ensure staff who use AI systems have an adequate level of AI literacy. A short, trackable training module with a certificate is the pragmatic way to evidence this, and it doubles as good change-management when rolling an assistant out to a workforce.</p>
<h2>Data governance and GDPR</h2>
<p>This is where on-premise deployment pays for itself in compliance terms:</p>
<ul><li>Data residency, when inference and documents stay on your hardware, residency is a fact of the architecture, not a clause to negotiate.</li><li>Right to erasure (GDPR Art. 17), erasure must propagate across conversations, documents, embeddings, any fine-tuned weights and backups; design for this from day one.</li><li>Lawful basis and minimisation, you still need a basis for processing and should apply DLP to keep sensitive categories out of prompts where appropriate.</li><li>Records of processing, your audit log is also your evidence base.</li></ul>
<h2>Logging and traceability</h2>
<p>High-risk systems must keep automatic logs, and even outside that tier an immutable, exportable audit trail is the difference between asserting compliance and demonstrating it. Look for SIEM export (syslog or signed webhook) so the AI system's records flow into the same monitoring as the rest of your estate.</p>
<h2>Human oversight</h2>
<p>The Act expects meaningful human oversight of consequential AI use. Architecturally this means role-gated capabilities, the ability to review and override, and not wiring an assistant directly into irreversible actions without a human in the loop.</p>
<h2>A condensed checklist</h2>
<ol><li>Classify each use case by risk tier and document it.</li><li>Add an Article 50 transparency notice and label AI output.</li><li>Roll out and track Article 4 AI-literacy training.</li><li>Pin data residency to your own hardware; confirm zero egress.</li><li>Implement erasure that propagates to embeddings, fine-tunes and backups.</li><li>Apply DLP to inbound and outbound messages.</li><li>Enable an immutable audit log with SIEM export.</li><li>Define human-oversight and role-gating policies.</li><li>Prepare DPIA/FRIA documentation where required.</li><li>Have EU-qualified counsel review the whole package.</li></ol>
<h2>Sector overlays</h2>
<p>On top of the horizontal obligations, expect sector rules: MDR/IVDR and EHDS in healthcare; DORA, CRA and EBA expectations in finance; FRIA, an Article 49 registry and NIS2 in the public sector; and national regimes such as Sweden's Dataskyddslagen and Offentlighetsprincipen. A platform that ships these overlays as templates rather than leaving you to build them is doing a meaningful amount of the work for you.</p>
<p>Want the downloadable version of this checklist, pre-filled with a risk catalogue and DPIA/FRIA templates? Get in touch and we will share the compliance kit.</p>]]></content:encoded>
    </item>
    <item>
      <title>Cloud AI vs. on-prem in 2026: a practical decision framework</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/cloud-ai-vs-onprem-2026-framework/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/cloud-ai-vs-onprem-2026-framework/</guid>
      <pubDate>Fri, 12 Jun 2026 09:00:00 GMT</pubDate>
      <category>Blog</category>
      <category>Field notes</category>
      <description>A four-question framework for deciding where your organisation's AI should run, data sensitivity, cost shape, dependency risk and operational reality.</description>
      <content:encoded><![CDATA[<p>“Should we run AI in the cloud or on our own hardware?” is usually asked as a technology question. It's really four smaller questions, about your data, your cost structure, your appetite for dependency, and your operational reality. Answer those honestly and the decision mostly makes itself.</p>
<h2>1. Can your most valuable data legally and culturally leave the building?</h2>
<p>Not “is the vendor's DPA acceptable”, can the data <em>leave</em>? For patient records, case files, deal documents, source code and citizen data, the answer is often no, regardless of what the contract says. If the data that would make AI most useful to you is exactly the data you can't send out, cloud AI gives you a capable assistant that isn't allowed to do its job.</p>
<div data-callout><p>A useful exercise: list the five document sets your team would most want an assistant to know. For each, ask whether you'd email it to an external consultancy. That's roughly the bar for sending it to any third party.</p></div>
<h2>2. Which cost shape fits, linear or flat?</h2>
<p>Per-seat cloud AI scales with headcount; owned hardware doesn't. Neither is universally better:</p>
<table><thead><tr><th>Factor</th><th>Cloud (per seat)</th><th>On-prem (owned)</th></tr></thead><tbody>
<tr><td>Cost shape</td><td>Linear with users</td><td>Flat after purchase</td></tr>
<tr><td>Small pilots</td><td>Cheap to start</td><td>Up-front hardware</td></tr>
<tr><td>Whole-company rollout</td><td>Compounds every month</td><td>Same box, more users</td></tr>
<tr><td>Budget character</td><td>Opex, recurring</td><td>Capex, owned asset</td></tr>
</tbody></table>
<p>The honest summary: for a five-person trial, cloud wins on cost. For an organisation-wide capability you intend to keep, the flat curve wins, usually sooner than people expect.</p>
<h2>3. How much dependency can your workflows tolerate?</h2>
<p>Cloud AI means your capability is downstream of someone else's roadmap: models get deprecated, terms get revised, prices move. That's survivable for casual use and painful for load-bearing workflows. The test we suggest: <em>if the vendor doubled the price or retired your model tomorrow, what would break?</em> If the answer is “our core processes”, you're not buying a tool, you're taking on a dependency, and dependencies deserve harder scrutiny.</p>
<h2>4. What can your team actually operate?</h2>
<p>The traditional argument for cloud was operational: someone else runs it. That argument has weakened as on-prem platforms became appliances. If your bar is “no Kubernetes, no ML engineers, one box, two cables”, that bar is now met. What remains true: you should evaluate on-prem the way you'd evaluate any infrastructure purchase, support, updates, monitoring, evidence for auditors, not as a science project.</p>
<h2>The shape of the answer</h2>
<p>In practice we see three stable outcomes. Teams with public data and spiky usage stay in the cloud. Regulated organisations with sensitive corpora go on-prem, because nothing else clears legal. And a growing middle group runs both: cloud for generic tasks, an appliance for everything touching the crown jewels. All three are rational, what's not rational is defaulting to the cloud without pricing in the data question, the cost curve, and the dependency.</p>
<p>If you're working through this for your own organisation, our <a href="/research/local-ai-sweden-eu-guide/">2026 guide to sovereign deployment in Sweden and the EU</a> goes deeper on the regulatory side, and <a href="/contact/">we're happy to talk through the maths</a> for your headcount.</p>]]></content:encoded>
    </item>
    <item>
      <title>Model Provenance and Supply-Chain Trust for Private LLMs</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/research/model-provenance-supply-chain-trust/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/research/model-provenance-supply-chain-trust/</guid>
      <pubDate>Thu, 11 Jun 2026 09:00:00 GMT</pubDate>
      <category>Research &amp; white papers</category>
      <category>Security</category>
      <description>Why model provenance matters for private LLM deployments, and how signing, hashing, pinning and trust scoring keep the weights you reviewed the weights you run.</description>
      <content:encoded><![CDATA[<p>When you download an open-weights model, you are adding a multi-gigabyte binary of opaque floating-point numbers to your most sensitive systems. Treated casually, that is one of the largest unreviewed dependencies in the enterprise. Model provenance is the discipline of knowing where those weights came from, that they have not changed, and whether they are fit to run, and it is becoming a procurement requirement for defence, public-sector and regulated buyers.</p>
<h2>Models are a supply-chain dependency</h2>
<p>We have spent a decade learning to scrutinise software supply chains, signing artefacts, pinning versions, generating SBOMs, scanning for tampering. Model weights have largely escaped that scrutiny despite carrying comparable risk: a poisoned or swapped checkpoint can subtly degrade outputs, leak training data, or behave differently from the artefact you evaluated. The fix is to apply the same supply-chain hygiene to models that we already apply to code.</p>
<h2>Hash on first load, then pin</h2>
<p>The foundational control is simple: compute a content hash of the weights the first time a model is loaded, record it, and pin to it. On every subsequent load the hash is checked. If a file changes, through tampering, corruption or an unnoticed substitution, it no longer matches and the system refuses to run it. This is the model equivalent of a checksummed, signed release, and it gives you a hard guarantee that the weights you reviewed are the weights serving your users.</p>
<blockquote><p>“Provenance is not about distrusting open models. It is about being able to prove, at any moment, that the thing running is exactly the thing you approved.”</p></blockquote>
<h2>A trust score, not a vibe</h2>
<p>Beyond integrity, administrators need to decide which models are appropriate to deploy at all. A structured provenance grade makes that decision auditable by scoring each model across dimensions such as:</p>
<ul><li>Source and license, is the origin reputable and the license compatible with your use?</li><li>Stability, is the model mature and widely exercised, or experimental?</li><li>Integrity, does it pass hash verification and load cleanly?</li><li>Behaviour, does it pass a canary evaluation on install before anyone relies on it?</li></ul>
<p>Rolling these into a single 0-100 grade, backed by a short canary eval at install time, turns 'which model should we trust?' from a matter of reputation into a documented, repeatable decision a security reviewer can sign off.</p>
<h2>Why it matters more on-premise</h2>
<p>Paradoxically, running your own models raises the provenance bar rather than lowering it. In a cloud service the provider implicitly owns model integrity; when you self-host, that responsibility moves to you. The upside is that you also gain the control to do it properly, pinning the exact weights, keeping them running regardless of upstream deprecation, and proving the chain of custody to an auditor. Provenance is the discipline that lets sovereignty and trust coexist.</p>
<p>Provenance is one plane of a larger sovereign-AI architecture. For the full picture, see our reference architecture for air-gapped LLM deployment and our primer on sovereign AI.</p>]]></content:encoded>
    </item>
    <item>
      <title>The European Health Data Space and AI in care: what changes</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/european-health-data-space-ai-in-care/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/european-health-data-space-ai-in-care/</guid>
      <pubDate>Wed, 10 Jun 2026 09:00:00 GMT</pubDate>
      <category>EU regulations</category>
      <category>Health data</category>
      <description>The EHDS Regulation entered into force in March 2025. What it means for health data, secure processing environments, and AI that runs where the data lives.</description>
      <content:encoded><![CDATA[<p>Health data is the most sensitive category most organisations will ever handle, and the EU has just built a whole legal framework around it. The <a href="https://eur-lex.europa.eu/eli/reg/2025/327/oj/eng">European Health Data Space Regulation</a>, Regulation (EU) 2025/327, entered into force on 26 March 2025 and rolls out in phases over the following years. If you work anywhere near clinical AI, it is worth understanding, because it says a lot about where the EU thinks sensitive data should be processed.</p>
<h2>Two uses, two sets of rules</h2>
<p>The EHDS splits health data use into two categories, and the distinction matters.</p>
<p>Primary use is what you would expect: using someone's health data to care for them. Electronic health records, cross-border access when a patient travels, that sort of thing. Secondary use is everything else, research, innovation, policymaking, training a model, and it runs through national health data access bodies that issue permits. The <a href="https://health.ec.europa.eu/ehealth-digital-health-and-care/european-health-data-space-regulation-ehds_en">Commission's EHDS pages</a> lay out the phased timeline, with the main secondary-use machinery arriving toward the end of the decade.</p>
<h2>The rule that should make AI vendors pay attention</h2>
<p>For secondary use, the EHDS does not just wave data out the door. Processing has to happen inside a secure processing environment that meets high privacy and cybersecurity standards. Data is pseudonymised or anonymised. And here is the striking part: you cannot download the personal data out of that environment. You bring your analysis to the data, run it inside the walls, and take only the results away.</p>
<p>Sit with that for a second, because it is a profound statement of principle. The most sensitive data in Europe, the EU has decided, should be processed in a controlled environment that the data does not leave. Researchers come to it. It does not go to them.</p>
<div data-callout><p>If that architecture sounds familiar, it should. "Bring the computation to the data, do not move the data to the computation" is exactly how on-premise AI works. The model runs where the records already live. Nothing is exported. The EHDS did not invent this pattern, but by writing it into law for health data it has blessed the idea that for truly sensitive information, local and sealed is the responsible default.</p></div>
<h2>Health data was always special, and now it has scaffolding</h2>
<p>Under <a href="https://eur-lex.europa.eu/eli/reg/2016/679/oj">GDPR Article 9</a>, health data is a special category with extra protection, and Swedish rules like the patient data act (patientdatalagen) add their own layer. The EHDS does not replace any of that. It builds infrastructure on top of it, and where AI is a medical device, the <a href="https://eur-lex.europa.eu/eli/reg/2017/745/oj">Medical Device Regulation</a> comes into play as well. It is a dense stack.</p>
<p>The practical takeaway for anyone building or buying clinical AI is not that the EHDS forces you on-premise today. The phased dates stretch out for years, and much of the detail lives in implementing acts still to come. The takeaway is about direction. When European regulators design their flagship health data framework around secure, no-download environments, they are telling you what "good" looks like for sensitive data. Running the AI where the data already sits is not a fringe position. It is increasingly the one the law is describing.</p>
<p>This is general information, not legal or clinical advice. Health data rules are among the most complex in the book, and the EHDS interacts with the GDPR, the MDR and national law in ways your compliance and clinical governance teams need to work through for your setting.</p>]]></content:encoded>
    </item>
    <item>
      <title>NIS2 and your AI infrastructure: what the directive means in practice</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/nis2-and-your-ai-infrastructure/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/nis2-and-your-ai-infrastructure/</guid>
      <pubDate>Fri, 05 Jun 2026 09:00:00 GMT</pubDate>
      <category>EU regulations</category>
      <category>NIS2</category>
      <description>NIS2 pulls tens of thousands of EU organisations into a real cybersecurity regime. If AI is becoming load-bearing infrastructure for you, it's in scope too.</description>
      <content:encoded><![CDATA[<p>NIS2 is the EU's broad cybersecurity directive, the successor to NIS1 with a much wider net. Member states were due to transpose it into national law by 17 October 2024, and it covers <em>essential</em> and <em>important</em> entities across sectors from energy, health and transport to digital infrastructure, manufacturing, and public administration. If your organisation is in scope, the security of the systems you run, increasingly including your AI systems, is now a regulated matter with management-level accountability.</p>
<h2>The obligations, in one paragraph</h2>
<p>In-scope entities must take “appropriate and proportionate” technical and organisational measures across a familiar checklist: risk analysis, incident handling, business continuity, <strong>supply-chain security</strong>, secure development and configuration, vulnerability handling, cryptography policies, access control and MFA. Add to that a strict incident-reporting clock (early warning within 24 hours of a significant incident, incident notification within 72) and personal responsibility for management bodies, and you have the shape of the regime. Penalties scale to €10M or 2% of global turnover for essential entities.</p>
<h2>Why your AI deployment is part of this</h2>
<p>Nothing in NIS2 says “artificial intelligence”. It doesn't need to. The directive regulates the network and information systems you rely on, and the moment an assistant is drafting your documents, answering questions over your files, or wired into your workflows via an API, it has become one of those systems. Three consequences follow:</p>
<ul>
<li><strong>It joins your risk analysis.</strong> What data can the AI system reach? Who can query it? What happens if it's compromised or unavailable? These now need documented answers.</li>
<li><strong>It joins your supply chain review.</strong> NIS2 makes you responsible for the security of your suppliers and service providers. Every external AI service is another supplier to assess, contract with, and monitor.</li>
<li><strong>It joins your incident scope.</strong> A breach involving the AI layer, leaked prompts, exfiltrated documents, a poisoned model, is an incident like any other, with the same 24/72-hour clocks.</li>
</ul>
<h2>The supply-chain angle favours shorter chains</h2>
<p>Here is the practical observation: every obligation above gets easier when the chain is shorter. A cloud AI dependency means your NIS2 supply-chain duties extend through the vendor and their sub-processors; your incident response depends on their disclosure speed; your continuity planning inherits their outages. A self-hosted deployment turns that external dependency into an internal system, still in scope, still needing hardening and monitoring, but assessed and evidenced by you, on your timeline.</p>
<div data-callout><p>This is not “on-prem is automatically compliant”, a badly run appliance is worse than a well-run cloud service. The point is narrower: NIS2 rewards architectures whose security you can inspect and evidence directly. Sealed perimeter, immutable audit log, SIEM export and controlled updates exist in Kaldryn precisely because that's what an auditor asks for.</p></div>
<h2>If you're starting now</h2>
<ol>
<li>Confirm whether you're in scope (sector + size, in your member state's transposition, in Sweden, via the new cybersecurity act implementing NIS2).</li>
<li>Add AI systems explicitly to your asset inventory and risk analysis, shadow AI included.</li>
<li>Map your AI supply chain: every external model API, plugin and integration is a supplier.</li>
<li>Rehearse the 24-hour early warning with an AI-flavoured scenario once, you'll find the gaps quickly.</li>
</ol>]]></content:encoded>
    </item>
    <item>
      <title>What an offline AI assistant can actually do</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/what-an-offline-ai-assistant-can-actually-do/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/what-an-offline-ai-assistant-can-actually-do/</guid>
      <pubDate>Thu, 28 May 2026 09:00:00 GMT</pubDate>
      <category>Blog</category>
      <category>Product</category>
      <description>People hear air-gapped and picture something crippled and slow. A model running with no internet is not a lesser assistant. Here is what it handles day to day.</description>
      <content:encoded><![CDATA[<p>Tell someone your AI assistant runs with no internet connection and watch their face. They picture something from a decade ago: slow, dim, barely able to string a sentence together. It is a stubborn misconception, and it is completely wrong. An assistant that runs entirely on local hardware, even fully air-gapped, is not a stripped-down version of the real thing. For day-to-day work it often is the real thing.</p>
<h2>Where the confusion comes from</h2>
<p>The mix-up is understandable. For years, "AI" meant "a website you log into," so no internet sounded like no AI. But the model itself does not need the internet to think. It needs the internet to be somewhere else, on a provider's servers. Move the model onto a machine in your building and the connection becomes optional. The intelligence was never in the network. It was in the weights, and the weights can sit right next to you.</p>
<h2>What it handles without ever phoning home</h2>
<p>Here is the ordinary, useful work an offline assistant does, none of which requires a connection to anything outside your walls. It answers questions about your own documents, because those documents live on the same network and the model reads them there. It drafts emails, reports, summaries and replies. It pulls structured information out of messy files: dates, names, figures, clauses. It classifies and routes. It helps people write code against your own systems. It holds a conversation over many turns. And it does all of this at full speed, because there is no round trip across the internet slowing it down.</p>
<div data-callout><p>The counterintuitive part: cutting the internet can make an assistant better for internal work, not worse. It is faster, because the answer does not travel to another country and back. It is more private, because there is nowhere for the data to go. And it keeps working during an internet outage that would take a cloud tool offline entirely. Unplug the cable and it does not even notice.</p></div>
<h2>What it genuinely cannot do</h2>
<p>Let me be honest about the trade, because the honest version is still a good deal. An offline assistant cannot look things up on the live web. It knows what is in your documents and what was in its training, not this morning's news or a public website. For a lot of internal work, that is exactly what you want, because the answers should come from your own trusted material, not a random web page. But if your use case genuinely needs live external information, an air-gapped setup is the wrong tool, and you can choose a connected configuration instead. The point is that it is your choice, not a limitation forced on you.</p>
<h2>The version most people actually want</h2>
<p>In practice, most organisations do not run fully air-gapped. They run connected for convenience and simply keep the AI workload sealed, so the sensitive processing never leaves the building even though the office has internet. The fully offline mode is there for the work that demands it: the classified project, the regulated data set, the environment where an internet connection is itself the risk. Either way, the assistant is capable, fast and private. "Offline" was never a synonym for "worse." It just means the intelligence came to your data instead of the other way round.</p>]]></content:encoded>
    </item>
    <item>
      <title>The real cost of on-premise AI versus per-seat cloud assistants</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/research/on-premise-ai-tco-versus-cloud/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/research/on-premise-ai-tco-versus-cloud/</guid>
      <pubDate>Wed, 20 May 2026 09:00:00 GMT</pubDate>
      <category>Research &amp; white papers</category>
      <category>Economics</category>
      <description>A grounded look at when running open models on your own hardware beats paying per seat or per token, and when it does not, with the break-even maths.</description>
      <content:encoded><![CDATA[<p>Every business case for local AI eventually runs into the same question from finance: is this actually cheaper than just paying a cloud provider? The honest answer is that it depends, and anyone who tells you otherwise is selling something. So let me lay out the real maths, including the parts that do not favour us.</p>
<h2>How cloud AI bills you</h2>
<p>There are two meters running. Per-seat products charge a flat monthly fee for each user. Per-token APIs charge for the text going in and, more expensively, the text coming out. As of 2026 a mid-tier frontier model runs somewhere around two to three dollars per million input tokens and ten to fifteen dollars per million output tokens, going by the published rates from <a href="https://platform.claude.com/docs/en/about-claude/pricing">Anthropic</a>, <a href="https://developers.openai.com/api/docs/pricing">OpenAI</a> and <a href="https://ai.google.dev/gemini-api/docs/pricing">Google</a>. There are also much cheaper small-model tiers, some down at ten to forty cents per million tokens, and for light workloads those are genuinely hard to beat.</p>
<p>That last point matters, so I will say it plainly. If a handful of people use AI occasionally, the cloud is probably cheaper than buying a GPU, and you should use the cloud. The economics of local AI do not turn on the simple case. They turn on volume.</p>
<h2>Where the meter starts to hurt</h2>
<p>Costs stop being trivial in three situations. The first is concurrency: lots of people using the assistant at once, all day. The second is retrieval over long documents, where every question drags a large chunk of context through the model and every token is billed. The third is agentic use, where the model calls itself in loops to complete a task, quietly multiplying the token count by ten or more. A single power user running agentic workflows can burn through hundreds of thousands of tokens a day. Multiply that by a department and the monthly invoice stops looking like a rounding error.</p>
<p>This is the profile where owning the hardware changes the picture, and it happens to be the exact profile of a company that wants AI woven into daily work rather than used as a novelty.</p>
<h2>The break-even, honestly</h2>
<p>There is a useful 2025 study on this, <a href="https://arxiv.org/html/2509.18101v3">a cost-benefit analysis of on-premise LLM deployment</a>, and its numbers are worth quoting because they are neither hype nor dismissal. A single modern consumer GPU can serve a 24 to 32 billion parameter model at roughly 150 to 200 tokens per second. For a small but steady workload, the authors find hardware paying for itself against commercial API costs in months, sometimes in under one. For larger deployments on data-centre GPUs, the payback stretches to a few years, and it only holds if you actually keep the hardware busy.</p>
<p>That caveat is the whole game. A GPU you bought and left idle is the most expensive way to run AI ever devised. The break-even assumes sustained utilisation, and independent analyses put the crossover point where self-hosting beats cloud at somewhere north of fifty percent steady usage. Below that, you are paying for silicon that naps.</p>
<div data-callout><p>The one-line summary for the finance meeting. Cloud AI is cheaper for light or spiky use. Owned hardware is cheaper for heavy, predictable use, and it pays back in months rather than years once the machine is genuinely busy. Pick the tool that matches your actual usage pattern, not the one with the better slide.</p></div>
<h2>The line item nobody budgets</h2>
<p>Here is where most do-it-yourself on-premise projects quietly bleed money. It is not the hardware. It is the person. Running your own inference stack means someone has to keep models updated, GPUs healthy, drivers current and security patched. Analyses of on-prem total cost of ownership repeatedly find that a fractional infrastructure engineer, call it half a full-time role at seventy-five to a hundred thousand a year, is the single largest three-year line item, larger than the servers.</p>
<p>This is the honest weak point of naive self-hosting, and it is also precisely the problem a managed appliance exists to solve. When the box arrives pre-configured and someone else handles the updates and monitoring, that staffing line disappears from your budget. The distinction between "run your own models" and "own an appliance that runs your models" is mostly this line item.</p>
<h2>The value that is not on the invoice</h2>
<p>Cost comparisons that stop at tokens miss the reason most of our customers are here in the first place. When the model runs in your building, your data never becomes a transfer to assess, a breach waiting to be reported, or a dependency on a provider that can change its prices or its terms. <a href="https://www.ibm.com/reports/data-breach">IBM's 2025 breach research</a> puts the global average cost of a data breach at 4.44 million dollars, and flags a real premium when unsanctioned AI tools are involved. None of that shows up in a per-token quote, but it is real money and real risk.</p>
<h2>How to actually decide</h2>
<p>Skip the vendor spreadsheets and do this instead. Estimate your steady-state usage in tokens per month, honestly, including the agentic and long-context workloads you plan to grow into. If it is small and likely to stay small, use the cloud and enjoy it. If it is large, or growing fast, or touching data you cannot send outside the building, price the owned option properly, including the staffing you will avoid with a managed appliance, and compare it to the cloud bill you would be signing up to pay every month, forever. For a lot of regulated, AI-hungry organisations, that comparison is not close. For others it genuinely is, and we would rather tell you so than win a deal you will regret.</p>]]></content:encoded>
    </item>
    <item>
      <title>Benchmarking open-weight models for regulated back-office work</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/research/benchmarking-open-models-regulated-work/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/research/benchmarking-open-models-regulated-work/</guid>
      <pubDate>Wed, 13 May 2026 09:00:00 GMT</pubDate>
      <category>Research &amp; white papers</category>
      <category>Architecture</category>
      <description>Open models are good enough for a lot of real enterprise work. A practical guide to proving that for your own tasks instead of trusting a leaderboard.</description>
      <content:encoded><![CDATA[<p>There is a stubborn myth that only a frontier cloud model can do useful work, and that anything you can run yourself is a toy. In 2026 that is simply out of date for the tasks most businesses actually care about. But the counter-myth, that open models are now just as good at everything, is also wrong. The truth is more useful than either slogan, and this paper is about finding it for your own work rather than arguing about leaderboards.</p>
<h2>Where open models are genuinely good now</h2>
<p>The enterprise sweet spot is drafting, summarising, classifying, extracting and answering questions over your own documents. On these tasks, leading open-weight models like Llama 4, Qwen 3.5, Mistral and DeepSeek score within a few points of closed frontier models on standard knowledge benchmarks, according to public leaderboards, and in practice are widely judged production-ready. If your use case is a private assistant that reads your files and helps people write, the model is not your bottleneck any more.</p>
<p>Where open models still trail is the hardest end: multi-step agentic coding, the trickiest reasoning, the frontier of the frontier. If that is your core workload, weigh it honestly. For the back office, it rarely is.</p>
<h2>Why leaderboards will mislead you</h2>
<p>Treat public benchmark scores with suspicion, for two reasons. First, the classic benchmarks are saturated. MMLU, the one everyone quotes, is so close to maxed out that a two-point difference between models is noise, not signal. The field has moved to harder tests of coding, reasoning and long-context retrieval precisely because the old ones stopped discriminating. Second, and more important, none of these tests are your work. A model that tops a leaderboard on general knowledge can still be mediocre at summarising your specific contracts in your specific house style. The only benchmark that predicts your results is your tasks.</p>
<div data-callout><p>The most valuable hour you can spend evaluating AI. Take twenty real examples of the work you actually want done: real documents to summarise, real questions to answer, real drafts to produce. Run them through the candidate models. Have the person who normally does that work grade the outputs. You will learn more in that hour than in a week of reading benchmark tables, because you measured the thing you care about instead of a proxy for it.</p></div>
<h2>Provenance and licence are part of the score</h2>
<p>For regulated work, raw capability is not the only axis. Two others matter as much. The first is provenance: where the weights came from, whether they are signed, whether you can pin the exact version you evaluated so it cannot change under you. The second is licence, which varies a lot between open models and decides whether you may actually use a given model commercially. It is worth noting that several of the strongest open models in 2026 come from Chinese labs, which is not a problem in itself when you run the weights entirely on your own hardware and nothing phones home, but it is a fact a careful buyer should know and reason about rather than discover later.</p>
<p>The reassuring part of running open models locally is that provenance becomes something you can actually control. You choose the model, you verify it, you host it, and it stays exactly as you left it. There is no silent upgrade that changes its behaviour mid-quarter, which is a real problem with rented models and a real headache for anyone who has to keep an evaluation valid over time.</p>
<h2>A short evaluation protocol</h2>
<p>Put together, a defensible model choice looks like this. Assemble a small suite of your real tasks with a known good answer for each. Shortlist two or three open models whose licence fits your use. Run the suite, blind if you can, and have a domain expert grade the outputs against your standard, not a generic one. Check provenance and pin versions. Re-run the suite when you change anything. It is not glamorous, and it is far more reliable than trusting a number someone else produced on a test that has nothing to do with your business.</p>
<p>Benchmark figures cited here come from public leaderboards rather than primary lab reports, and model capabilities move quickly. Treat this as a method, not a ranking. The method outlives the models.</p>]]></content:encoded>
    </item>
    <item>
      <title>How to run an AI pilot your security team will approve</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/how-to-run-an-ai-pilot-your-security-team-will-approve/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/how-to-run-an-ai-pilot-your-security-team-will-approve/</guid>
      <pubDate>Fri, 08 May 2026 09:00:00 GMT</pubDate>
      <category>Blog</category>
      <category>AI in practice</category>
      <description>Most AI pilots die in the security review, and usually for avoidable reasons. A practical guide to designing one that gets a yes instead of a maybe.</description>
      <content:encoded><![CDATA[<p>Plenty of AI pilots never make it out of the security review, and it is rarely because the technology failed. It is because the pilot was designed to impress the business and then handed to the security team as an afterthought, with all the awkward questions unanswered. If you want a yes, design for the review from day one. Here is how.</p>
<h2>Answer the one question they will ask first</h2>
<p>Every security review of an AI tool starts, explicitly or not, with a single question: where does our data go? If your answer involves a third-party provider, another country, or a sentence with the word "trust" in it, you have just created a week of follow-up questions about transfers, retention, sub-processors and breach exposure. If your answer is "it stays on our own hardware, on our own network, and here is how we prove it," you have skipped most of the interrogation. Design the pilot so the answer is the short one.</p>
<h2>Scope it small and specific</h2>
<p>Vague pilots frighten security teams because vague pilots cannot be assessed. "We will roll out AI across the company" is not a plan, it is a liability. "We will let the support team summarise tickets, using this data set, on this system, for eight weeks, with these people having access" is something a reviewer can actually evaluate and approve. Pick one workflow, one data set, one group of users. Make the boundaries obvious.</p>
<div data-callout><p>A pilot your security team can reason about beats an impressive one they cannot. Give them a clear scope, a clear data flow and a clear off switch, and you have handed them the tools to say yes. Give them ambition and hand-waving, and you have handed them every reason to say not yet.</p></div>
<h2>Bring the paperwork they were going to ask for</h2>
<p>You can save weeks by showing up with the documents the review needs instead of waiting to be asked. For anything touching personal data, a data protection impact assessment is often required under <a href="https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng">Article 35 of the GDPR</a>, and a pilot that keeps data in-house makes that document short, because there is little transfer risk to assess. Have the access list ready. Have the logging story ready, so you can show what the system did and who used it. Know which regulatory buckets your use falls into, including the <a href="https://artificialintelligenceact.eu/article/4/">AI literacy duty</a> that already applies to staff using AI. None of this is heavy if you plan for it. All of it is painful if you improvise it under questioning.</p>
<h2>Give them an off switch and an audit trail</h2>
<p>Two things make a reviewer relax more than almost anything else. The first is a clean way to stop: if the pilot goes wrong, how fast can you shut it down and contain it? Have an answer. The second is an audit trail: can you show, after the fact, exactly what the system did, with which data, for whom? A system that keeps an immutable log turns "trust us" into "look for yourself," and that shift is often the difference between approval and delay.</p>
<p>None of this is about gaming the review. It is about respecting it. Security teams are not the enemy of an AI rollout. They are the reason the rollout survives contact with the real world. Design the pilot as though the review is the point, keep the data in the building so the hardest questions answer themselves, and you will spend far less time defending it and far more time using it.</p>]]></content:encoded>
    </item>
    <item>
      <title>Article 50 in practice: telling people when they are talking to AI</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/ai-act-article-50-transparency-in-practice/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/ai-act-article-50-transparency-in-practice/</guid>
      <pubDate>Wed, 06 May 2026 09:00:00 GMT</pubDate>
      <category>EU regulations</category>
      <category>EU AI Act</category>
      <description>From August 2026 the AI Act requires you to disclose AI interactions and label AI-generated content. A practical look at what Article 50 asks for.</description>
      <content:encoded><![CDATA[<p>Not every part of the AI Act is about high-risk systems and conformity assessments. <a href="https://artificialintelligenceact.eu/article/50/">Article 50</a> is the one that touches almost everybody, because it is about something simple: honesty about when a person is dealing with a machine. It applies from 2 August 2026, and unlike the high-risk deadlines, it did not move in the 2026 reshuffle.</p>
<h2>The four things it asks for</h2>
<p>Article 50 sets out transparency duties for a handful of situations. In plain terms:</p>
<ul>
<li>If people interact with an AI system such as a chatbot, they have to be told, unless it is obvious to a reasonably observant person.</li>
<li>If a system generates synthetic audio, image, video or text, that output has to be marked in a machine-readable way as artificially generated.</li>
<li>If you use emotion recognition or biometric categorisation on people, you have to inform them.</li>
<li>If you produce a deepfake, or AI-generated text published to inform the public on matters of public interest, you have to disclose that it is artificial, with some exceptions for artistic or clearly editorial work.</li>
</ul>
<p>The Commission's <a href="https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai">regulatory framework pages</a> are the place to watch for the detailed guidance, which tends to arrive close to the application date.</p>
<h2>Why the machine-readable bit is harder than it sounds</h2>
<p>Telling a user "you are chatting with a bot" is easy. Marking generated content in a machine-readable way is the part that trips teams up. It means watermarking or metadata that survives copying, screenshots and re-encoding, which is a genuinely unsolved problem for text and a partly solved one for images. The law is asking for something the technology does not do perfectly yet, and everyone involved knows it.</p>
<p>For most companies deploying an internal assistant, the practical exposure is narrow. You are not usually generating public deepfakes. You are running a tool that drafts emails and answers questions over your own files. The relevant duty is mostly the first one: make sure people know when they are talking to the assistant rather than a colleague, which is easy to design in.</p>
<div data-callout><p>A useful test. Walk through every place your AI touches a human who is not the operator. A customer on a support chat. A citizen filling in a form. An employee reading an AI-drafted notice. For each, ask whether that person can tell a machine was involved. If the answer is no and it matters, that is where Article 50 lives.</p></div>
<h2>Where running things locally helps, and where it does not</h2>
<p>Let me be honest about the limits here. Running a model on your own hardware does not exempt you from Article 50. The duty follows the use, not the location of the GPU. If you generate a deepfake on a local box, you still have to label it.</p>
<p>What local deployment does change is your ability to build the disclosure in and keep it there. When you control the whole stack, you can enforce a labelling policy at the point of generation, log every output, and prove after the fact what the system produced and how it was marked. When the generation happens behind a third-party API, you are trusting a supplier's implementation and their roadmap. For a duty that turns on being able to demonstrate what your system did, owning the system is a quieter life.</p>
<p>This briefing is general information and not legal advice. Article 50 has carve-outs and edge cases, particularly around journalism and creative work, so check the final text and any Commission guidance against your specific use with a qualified adviser.</p>]]></content:encoded>
    </item>
    <item>
      <title>The security case for keeping AI inference on-premise</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/research/security-case-for-on-premise-ai-inference/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/research/security-case-for-on-premise-ai-inference/</guid>
      <pubDate>Wed, 29 Apr 2026 09:00:00 GMT</pubDate>
      <category>Research &amp; white papers</category>
      <category>Security</category>
      <description>A threat model for enterprise AI: what actually goes wrong, why shadow AI is a breach vector, and how keeping inference in-house shrinks the attack surface.</description>
      <content:encoded><![CDATA[<p>Security conversations about AI tend to get lost in exotic scenarios: prompt injection, model poisoning, jailbreaks. Those are real and worth attention. But for most organisations, the largest risk is duller and more immediate. It is that sensitive data walks out the door inside ordinary AI usage, and nobody logged it leaving. This paper builds the threat model from that starting point.</p>
<h2>The threat that is already happening</h2>
<p>Before anyone attacks your model, your own people are quietly moving data into tools you do not control. It is called shadow AI, and it is not hypothetical. <a href="https://www.ibm.com/reports/data-breach">IBM's 2025 Cost of a Data Breach research</a> found that one in five breached organisations were breached in connection with shadow AI, that these incidents carried a cost premium of hundreds of thousands of dollars, and that the great majority of organisations that suffered an AI-related breach lacked proper access controls on those AI tools. The global average breach already sits at 4.44 million dollars. Shadow AI is not a compliance footnote. It is a live breach vector with a price tag.</p>
<p>The mechanism is mundane. An employee pastes a customer list, a contract, a patient summary or source code into a public assistant to get help. The data is now on someone else's servers, possibly retained, possibly used to improve a model, and entirely outside your logs and your control. No firewall stopped it because it left through a browser tab.</p>
<h2>The supply-chain dimension</h2>
<p>Every external AI service you adopt is another supplier in a chain that regulators increasingly expect you to secure. <a href="https://eur-lex.europa.eu/eli/dir/2022/2555/oj">NIS2</a> made supply-chain security a statutory obligation for in-scope entities. <a href="https://eur-lex.europa.eu/eli/reg/2022/2554/oj">DORA</a> created direct oversight of critical cloud providers precisely because so much of European finance leans on the same few of them. The EU cybersecurity agency <a href="https://www.enisa.europa.eu/publications/threat-landscape-for-supply-chain-attacks">ENISA has documented</a> how supply-chain attacks reach their targets not head-on but through a trusted supplier. Each AI dependency you add is one more link that can be attacked, and one more relationship you have to assess, contract for and monitor.</p>
<div data-callout><p>A useful reframing for a security review. Do not ask only "is this AI vendor secure?" Ask "what new attack surface and what new supplier does this AI create, and could we avoid creating it at all?" Sometimes the most secure integration is the one you did not build, because the capability lived inside your own boundary instead.</p></div>
<h2>What on-premise inference removes</h2>
<p>Keeping the model inside your network does not make you invulnerable. You still have to secure the box, patch it, control access to it and watch it. What it does is remove entire categories of risk rather than adding controls to manage them.</p>
<p>When inference happens in your building, the exfiltration path through a third-party API is closed, because there is no third-party API in the sensitive path. The supplier you would have to secure and monitor is not in the chain. The international transfer you would have to justify never occurs. The provider that could log your prompts, change its retention policy, or be compelled to hand data to a foreign authority is not holding your data, because your data never reached it. A sealed system with an immutable audit log lets you prove, rather than hope, that sensitive information stayed where it belonged.</p>
<h2>Building the model</h2>
<p>Put the pieces together and the threat model for enterprise AI has a clear centre of gravity. The headline risk is not a clever adversary defeating your model. It is routine data leakage through unsanctioned or externally hosted tools, compounded by a growing web of AI suppliers you now have to secure. The most effective single control is not another monitoring product. It is architecture: give people a capable assistant inside the boundary so they stop reaching for the ones outside it, and keep the sensitive inference on hardware you own so the data has nowhere external to go. Everything else, the prompt-injection defences and the rest, is worth doing on top of that foundation. It is not a substitute for it.</p>
<p>This paper is general guidance, not a security assessment of your environment. Your own threat model depends on your data, your sector and your obligations, and is a conversation for your security team. The direction it points, though, is consistent with where European regulation has been heading for years. Shorten the chain. Keep the sensitive data home.</p>]]></content:encoded>
    </item>
    <item>
      <title>Own it or rent it: two ways to buy AI</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/own-it-or-rent-it-two-ways-to-buy-ai/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/own-it-or-rent-it-two-ways-to-buy-ai/</guid>
      <pubDate>Wed, 22 Apr 2026 09:00:00 GMT</pubDate>
      <category>Blog</category>
      <category>Field notes</category>
      <description>Underneath every AI procurement is a single old question that has nothing to do with technology: do you want to own this capability, or rent it forever?</description>
      <content:encoded><![CDATA[<p>Strip away the model names, the benchmarks and the demos, and every decision about how to buy AI comes down to a question people have been asking about everything from tools to houses for a very long time. Do you want to own it, or rent it? Neither answer is wrong. But a lot of companies drift into renting without ever noticing they made a choice, and that is worth avoiding.</p>
<h2>What renting gives you</h2>
<p>Renting AI, which is what using a cloud provider mostly is, has real advantages and I am not going to pretend otherwise. You start immediately. You pay nothing up front. Someone else runs the infrastructure, patches it and keeps it healthy. You always have access to the latest model the moment it ships. For getting going, for experiments, for light and occasional use, renting is genuinely the sensible choice. If that is your situation, rent, and do not let anyone guilt you out of it.</p>
<h2>What renting costs you</h2>
<p>The catch with renting is the catch with all renting. You pay forever, the payments scale with how much you use, and at the end you own nothing. Your data lives in the landlord's building, under the landlord's rules, which can change. You are exposed to price rises and to a model being retired out from under a workflow you built on it. And the more essential the capability becomes to how you work, the more your business depends on a thing you do not control. Dependence is the real rent, and it compounds quietly.</p>
<div data-callout><p>A simple test for which mode you are in. If the AI vendor doubled the price tomorrow, or retired the model your team relies on, or was ordered to hand over its logs, what would happen to you? If the answer is "not much," you are comfortably renting. If the answer is "our core processes would break," you are not really renting a tool any more. You are depending on one, and dependence deserves a harder look.</p></div>
<h2>What owning gives you</h2>
<p>Owning AI means running it on your own hardware. It costs more to start and it asks more of you, or of whoever manages the box for you. In exchange you get the things ownership always gives. A fixed cost instead of a meter. Data that stays in your building. A model you can pin so it does not change under you. Independence from a provider's prices, terms and roadmap. For heavy, steady, sensitive use, the sums and the control both tend to favour owning, which is the case we make in detail in our note on <a href="/research/on-premise-ai-tco-versus-cloud/">the total cost of ownership of on-premise AI</a>.</p>
<h2>Just make it a decision</h2>
<p>The point of this is not that owning always wins. It does not. Plenty of companies should rent, and plenty should run both, cloud for the casual and generic, owned for the crown jewels. The point is to choose on purpose. Renting AI is easy to slide into and hard to climb out of once your workflows are built on someone else's meter. Before that happens, spend an hour on the old question. For this capability, at our scale, with our data, do we want to own it or rent it? Answer it deliberately, and you will make a better choice than the one you would have drifted into.</p>]]></content:encoded>
    </item>
    <item>
      <title>Schrems II, six years on: where EU to US data transfers stand in 2026</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/schrems-ii-eu-us-data-transfers-2026/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/schrems-ii-eu-us-data-transfers-2026/</guid>
      <pubDate>Wed, 15 Apr 2026 09:00:00 GMT</pubDate>
      <category>EU regulations</category>
      <category>Data transfers</category>
      <description>The Schrems II ruling still shapes every EU to US data transfer. A look at the Data Privacy Framework, its wobble, and why some data should just stay home.</description>
      <content:encoded><![CDATA[<p>Six years ago, on 16 July 2020, the Court of Justice of the European Union handed down the ruling everyone calls <a href="https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:62018CJ0311">Schrems II</a>. It struck down the Privacy Shield, the arrangement that had let companies move personal data from the EU to the United States, and it did so because US surveillance law gave European citizens no real way to object. That decision still shapes every transatlantic data flow, including the ones hiding inside your AI tools.</p>
<h2>A quick recap for people who tuned out</h2>
<p>The court's problem was never really with any single company. It was with the mismatch between EU data protection law and US national security law. Programmes operating under authorities like FISA Section 702 let US agencies reach data held by US providers, and a European whose data got swept up had no equivalent of a court to complain to. That gap is why Privacy Shield fell, and why its predecessor Safe Harbor fell before it.</p>
<p>After Schrems II, companies leaned on Standard Contractual Clauses plus so-called supplementary measures, and everyone had to run a transfer impact assessment to check whether the destination country actually offered adequate protection. It was a lot of paperwork for something that, deep down, the paperwork could not fully fix.</p>
<h2>The Data Privacy Framework, and why people are nervous</h2>
<p>In July 2023 the Commission adopted a new <a href="https://commission.europa.eu/law/law-topic/data-protection/international-dimension-data-protection/eu-us-data-transfers_en">adequacy decision</a>, the EU to US Data Privacy Framework, which restored a legal basis for transfers to certified US companies. Useful, and a relief for a lot of businesses. But it sits on the same fault line as the two arrangements that came before it, and the same activist who brought down the first two has said he expects this one to end up back in Luxembourg.</p>
<p>Through 2025 and into 2026 the framework has been under real pressure, including questions about the independence of the US oversight body it relies on. Nobody sensible is treating it as settled. If your compliance plan assumes the Data Privacy Framework will still be standing in three years, you are making a bet, not a plan.</p>
<div data-callout><p>Here is the thing about transfer risk. Every mechanism in this area, Privacy Shield, SCCs, the Data Privacy Framework, is a way of managing the consequences of your data leaving the EU. None of them changes the underlying fact that it left. The only move that removes the risk rather than managing it is not transferring the data at all.</p></div>
<h2>Where AI quietly reintroduces the problem</h2>
<p>Most companies spent years getting their transfer house in order, and then rolled out AI tools that undo a lot of it in a single procurement. A cloud assistant sends prompts, documents and context to a provider that is very often US-controlled, which pulls you straight back into the Schrems II world of adequacy decisions and transfer assessments. The AI did not create a new legal category. It just poured new data into the old one.</p>
<p>This is the least glamorous argument for running models locally, and maybe the strongest. When the inference happens on a server in your building, there is no transfer to assess, no adequacy decision to depend on, and no activist lawsuit that can change your legal exposure overnight. The data that would have crossed the Atlantic never leaves the room. The <a href="https://www.edpb.europa.eu/">European Data Protection Board</a> keeps publishing guidance on transfers, and it is worth reading, but the shortest compliant path is the one where the guidance does not apply to your workload because nothing crossed a border.</p>
<p>None of this is legal advice. International transfer rules are intricate and moving, and your DPO and counsel should map your specific flows. The general point stands on its own though. Data that stays home is data you never have to defend at the border.</p>]]></content:encoded>
    </item>
    <item>
      <title>SOC 2 and ISO 27001 for a self-hosted AI platform: mapping the controls</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/research/soc2-iso27001-self-hosted-ai-controls/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/research/soc2-iso27001-self-hosted-ai-controls/</guid>
      <pubDate>Wed, 08 Apr 2026 09:00:00 GMT</pubDate>
      <category>Research &amp; white papers</category>
      <category>Compliance</category>
      <description>What SOC 2 and ISO 27001 actually certify, how they differ, and how a self-hosted AI platform maps onto their controls instead of inheriting a vendor's.</description>
      <content:encoded><![CDATA[<p>Ask a security team what they need before they will approve an AI tool and two acronyms come up fast: SOC 2 and ISO 27001. They get treated as interchangeable trust badges, which is a mistake, because they are different instruments that answer different questions. If you run AI in-house, understanding the difference tells you what evidence you actually have to produce, and where owning the system helps.</p>
<h2>What SOC 2 actually is</h2>
<p>SOC 2 is an attestation. An independent auditor examines your controls against the <a href="https://www.aicpa-cima.com/resources/download/2017-trust-services-criteria-with-revised-points-of-focus-2022">Trust Services Criteria</a> published by the American Institute of CPAs and writes an opinion on them. There are five criteria: Security, Availability, Processing Integrity, Confidentiality and Privacy. Only Security, sometimes called the common criteria, is mandatory. You choose the others based on what you promise customers.</p>
<p>There are also two flavours. A Type I report looks at whether your controls are well designed at a single point in time. A Type II report watches whether they actually operated effectively over a period, usually six to twelve months. Enterprise buyers who know what they are doing ask for Type II, because a snapshot of good intentions is easy and a track record is not.</p>
<h2>What ISO 27001 actually is</h2>
<p>ISO 27001 is a certification, which is a different animal from an attestation. Instead of an auditor's narrative opinion, you get a pass-or-fail certificate from an accredited body that says your information security management system meets the standard. The 2022 edition organises 93 controls into four themes: organisational, people, physical and technological. You do not implement all 93 blindly. You run a risk assessment, decide which apply, and record your reasoning in a statement of applicability.</p>
<p>The useful news is that these two frameworks are not rivals so much as cousins. The overlap between the SOC 2 criteria and the ISO 27001 controls is large, commonly estimated around eighty percent, so the evidence you gather for one carries most of the weight for the other. Doing both is far less than twice the work.</p>
<div data-callout><p>The distinction that matters when a vendor waves a certificate at you. A cloud AI provider's SOC 2 report attests to the provider's controls, in the provider's environment. It says almost nothing about your deployment, your data handling, or the specific way you use the tool. Their badge is not your compliance. When you run the platform yourself, the controls being audited are yours, which is more work and also more honest.</p></div>
<h2>Where self-hosting maps cleanly</h2>
<p>Several of the criteria and controls land naturally when the system lives in your own environment. Confidentiality, one of the five Trust Services Criteria, is straightforward to evidence when the data never leaves a boundary you control and the encryption keys are held by you rather than a third party. Access control, a whole cluster of ISO 27001 technological controls, is yours to enforce end to end. Physical security, an ISO theme in its own right, becomes a question about your own premises rather than a paragraph of trust in someone else's data centre you will never visit.</p>
<p>The flip side is real and worth stating. Owning the controls means owning the evidence. Nobody hands you a report to forward. You have to run the access reviews, keep the logs, test the backups and document the lot. A well-built self-hosted platform makes that easier by generating the audit trail for you, but the responsibility is yours. That is the trade: less borrowed trust, more of your own.</p>
<h2>How to think about it as a buyer or a builder</h2>
<p>If you are evaluating AI, do not accept a vendor's SOC 2 as evidence about your own use of their tool. Ask what controls you are responsible for and whether the architecture lets you satisfy them. If you are building an internal capability, decide early which criteria and controls you will need to show, because retrofitting an audit trail onto a system that was not designed to keep one is miserable. Either way, the frameworks are not the goal. They are a structured way of proving that the sensible things are actually being done, and a self-hosted design gives you direct control over most of the sensible things.</p>
<p>This paper is general information rather than a compliance opinion. The exact scope of a SOC 2 examination or an ISO 27001 certification for your organisation is a conversation for your auditor and your security function.</p>]]></content:encoded>
    </item>
    <item>
      <title>You don't need a frontier model to read your own documents</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/you-dont-need-a-frontier-model-for-your-documents/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/you-dont-need-a-frontier-model-for-your-documents/</guid>
      <pubDate>Thu, 02 Apr 2026 09:00:00 GMT</pubDate>
      <category>Blog</category>
      <category>AI in practice</category>
      <description>The best model in the world is wasted on most business tasks. For drafting and answering questions over your files, open models you can run yourself are enough.</description>
      <content:encoded><![CDATA[<p>A lot of AI strategy is quietly built on a status assumption: that you need the biggest, newest, most talked-about model, and anything less is settling. For a narrow set of genuinely hard problems, that is true. For the work most businesses actually want done, it is not even close to true, and the assumption is costing people money and control.</p>
<h2>Match the model to the job</h2>
<p>Think about what an internal assistant is really asked to do. Summarise this document. Answer a question using our own files. Draft a reply in our house style. Pull the key dates out of this contract. Classify these tickets. This is the bread and butter of business AI, and none of it is a frontier reasoning challenge. It is competent language work over your own material.</p>
<p>For that work, the open-weight models you can run on your own hardware in 2026 are genuinely good. Models like Llama 4, Qwen, Mistral and others land within a few points of the big closed models on standard knowledge benchmarks, and in practice they are perfectly capable of the drafting, summarising and document questions that make up the day job. Using a frontier cloud model for this is like chartering a jet to cross the street. It works, but you have badly overpaid, and you handed someone else your luggage.</p>
<h2>Where the big models still earn their keep</h2>
<p>Let me be fair, because the honest version of this argument is more persuasive than the hype version. There are tasks where frontier models still pull ahead: the hardest multi-step reasoning, complex agentic coding, the genuine edge of what is possible. If that is your core business, weigh it seriously and pay for the best. But for a company whose AI need is a private assistant over its own knowledge, that gap is irrelevant. You are paying a premium, in money and in data exposure, for capability you will never use on the task in front of you.</p>
<div data-callout><p>The test that settles it is not a benchmark. Take twenty real examples of the work you want done, run them through a model you could actually host, and have the person who normally does that work grade the results. Most teams are surprised how good "good enough" turns out to be, and how little the leaderboard had to do with it. We wrote up the full method in our piece on <a href="/research/benchmarking-open-models-regulated-work/">benchmarking open models for regulated work</a>.</p></div>
<h2>What "good enough" buys you</h2>
<p>Once you accept that an open model handles the real work, the rest of the decision opens up. You can run it yourself, which means your documents never leave the building. You can pin the exact version so it does not change under you mid-quarter, which anyone maintaining a rented model has learned to dread. You pay a fixed cost for hardware instead of a meter that runs forever. And you stop routing your most sensitive internal knowledge through a third party for tasks that never needed one.</p>
<p>The frontier is a wonderful thing, and for the problems that need it, nothing else will do. Just notice how few of your actual problems live there. Most of them are waiting quietly in your own files, and a model you can hold in your own hands is more than enough to answer them.</p>]]></content:encoded>
    </item>
    <item>
      <title>Data protection by design for LLM deployments: a technical white paper</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/research/data-protection-by-design-for-llm-deployments/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/research/data-protection-by-design-for-llm-deployments/</guid>
      <pubDate>Wed, 25 Mar 2026 09:00:00 GMT</pubDate>
      <category>Research &amp; white papers</category>
      <category>Compliance</category>
      <description>How GDPR principles like minimisation, erasure and lawful basis actually land on a working AI system, and why architecture decides whether you can comply.</description>
      <content:encoded><![CDATA[<p>Article 25 of the <a href="https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng">GDPR</a> asks for data protection "by design and by default." It is one of those phrases that sounds like a poster in an office kitchen until you try to build a real AI system and discover that the poster was describing a set of engineering decisions you already got wrong. This paper is about making those decisions on purpose.</p>
<h2>The principles that bite hardest</h2>
<p>A working language-model deployment collides with a handful of GDPR principles more than the rest. Purpose limitation and data minimisation, from Article 5, are in constant tension with the machine-learning instinct to hoard data because it might be useful. Lawful basis, from Article 6, has to be settled before you process a single record, and for most internal tools that means a documented legitimate-interest assessment. Special-category data, from Article 9, is the landmine: health, biometric, political and similar data is prohibited to process unless a specific exception applies, and it has a way of turning up in document sets nobody screened.</p>
<p>None of this is exotic. It is the ordinary GDPR you already know, applied to a system that happens to read everything you feed it and remember more than you expect.</p>
<h2>The erasure problem is real, and most vendors are vague about it</h2>
<p>Article 17, the right to be forgotten, is where AI architecture stops being a formality. Deleting a person's record from a database is a solved problem. Deleting their data from a model that was trained on it, or from the vector embeddings a retrieval system built out of their documents, is not. The information is smeared across weights and numbers in a way you cannot simply drop a row from.</p>
<p>When you ask any AI vendor how erasure works, listen for specifics. What happens to the embeddings derived from a deleted document? How long until it ages out of backups? Does a fine-tuned model retain it? If the answer is a shrug or a slogan, the erasure story is not real. The <a href="https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf">EDPB's Opinion 28/2024</a> makes clear this is not a theoretical worry: the Board stated that a model trained on personal data cannot, in every case, be assumed anonymous, and that anonymity has to be assessed rather than declared.</p>
<div data-callout><p>A practical erasure test for any AI system you are evaluating. Delete a document, then ask the system a question that only that document could answer. If it still answers, the data was not really erased, it was merely hidden from the file list. A defensible design propagates deletion through conversations, documents, embeddings, any fine-tunes and backups, and logs that it did.</p></div>
<h2>Why architecture, not paperwork, decides this</h2>
<p>You can write the finest data protection impact assessment in the world, as Article 35 often requires for high-risk AI, and it will not save you if the underlying system cannot honour a deletion request or keeps sending personal data somewhere it should not go. The document describes the system. It does not fix it.</p>
<p>This is the quiet argument for building the boundary in early. When the model, the index and the logs all sit inside your own network, the mitigations section of your DPIA gets dramatically shorter, because whole categories of risk are absent rather than managed. No third-party processor for the inference step. No international transfer to assess against the shifting backdrop of adequacy decisions. No sub-processor chain that changes when a vendor signs a new supplier. The <a href="https://www.edpb.europa.eu/our-work-tools/our-documents/opinion-board-art-64/opinion-282024-certain-data-protection-aspects_en">EDPB's guidance</a> spends pages on how to handle these risks in a cloud setting. Removing the cloud from the sensitive path removes the pages.</p>
<h2>A design checklist worth keeping</h2>
<p>If you take one thing from this paper, make it a set of questions to ask at design time rather than audit time. Where does personal data physically go when a prompt is submitted? Who, human or corporate, can technically read it along the way? Can a deletion request be honoured completely and demonstrably? Is special-category data filtered before it reaches the model? Is there an immutable log of what the system did with whom? A deployment that can answer those cleanly is one where by-design was more than a slogan.</p>
<p>This paper is general information, not legal advice. The GDPR interacts with the AI Act, sector law and national rules in ways your data protection officer and counsel need to map for your circumstances. The engineering point stands on its own, though. Compliance you can prove starts with architecture you chose deliberately.</p>]]></content:encoded>
    </item>
    <item>
      <title>Electricity, not tokens: budgeting for local AI</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/electricity-not-tokens-budgeting-for-local-ai/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/electricity-not-tokens-budgeting-for-local-ai/</guid>
      <pubDate>Thu, 19 Mar 2026 09:00:00 GMT</pubDate>
      <category>Blog</category>
      <category>Economics</category>
      <description>When you own the hardware, the running cost of AI stops being a per-question meter and turns into something much older and steadier: an electricity bill.</description>
      <content:encoded><![CDATA[<p>There is a mental shift that happens when a company moves from cloud AI to running its own. The cost of using the thing stops being a meter that ticks on every question and turns into something much more familiar and much less stressful: a power bill. It is worth sitting with that change, because it reshapes how you budget.</p>
<h2>From marginal cost to fixed cost</h2>
<p>With per-token cloud AI, every use has a price. Ask more questions, pay more. The cost scales with your success. When you own the hardware, that relationship breaks. The machine costs what it costs whether it answers ten questions a day or ten thousand. Past the purchase, the main thing you are paying for is the electricity to run it and cool it. Usage is effectively free. The marginal cost of one more question rounds to nothing.</p>
<p>This flips the psychology of AI adoption in a way people underestimate. On a meter, you quietly discourage use, because use is cost. On owned hardware, you want people to use it as much as possible, because you have already paid for it and idle capacity is wasted money. The incentives finally point the same way as the goal.</p>
<h2>What the electricity actually is</h2>
<p>Let me not hand-wave the number, because energy is a real line item. A serious inference server with high-end GPUs draws real power, and cooling adds a meaningful amount on top. Independent cost analyses in 2026 put a single high-end GPU server in the region of ten kilowatts under load, with cooling adding perhaps a quarter to a third again. That is a genuine running cost, and any honest budget for local AI includes it rather than pretending inference is free.</p>
<div data-callout><p>The break-even question is really an electricity question in disguise. Cloud AI charges you per token forever. Owned hardware charges you the up-front cost plus power. At high, steady usage, the power bill is far smaller than the token bill would have been, which is why heavy users come out ahead. At light usage, the hardware sits half-idle burning electricity for little benefit, and the cloud wins. Your usage pattern decides which story is yours.</p></div>
<h2>The part that is easy to like</h2>
<p>There is a quieter benefit to the electricity model beyond the maths. A power bill is predictable. You can budget it a year out and it will not surprise you because a team discovered a new agentic workflow and tripled the token count overnight. Owned hardware gives you a cost you can plan around, which finance departments tend to appreciate more than a variable bill that grows with adoption.</p>
<p>And for a company running in a place with clean, steady, reasonably priced power, the electricity model has an environmental and a cost logic that both point the same way. We build in Sweden, where the grid is largely low-carbon, and running your own inference on that kind of power is a different proposition than doing it on a dirtier, pricier grid. The precise figures depend on your location and your tariff, so price it for your own situation rather than trusting a brochure. But the shape of the deal is appealing: pay once for the machine, then pay for power, and stop feeding a meter that never stops running.</p>]]></content:encoded>
    </item>
    <item>
      <title>AI literacy is now a legal duty: reading Article 4 of the AI Act</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/ai-literacy-article-4-now-the-law/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/ai-literacy-article-4-now-the-law/</guid>
      <pubDate>Wed, 11 Mar 2026 09:00:00 GMT</pubDate>
      <category>EU regulations</category>
      <category>EU AI Act</category>
      <description>Since February 2025 the AI Act has required organisations to make sure staff who use AI understand it. What Article 4 asks for, and how to satisfy it.</description>
      <content:encoded><![CDATA[<p>There is an obligation in the AI Act that has been in force since 2 February 2025 and that a surprising number of organisations have not clocked. It is short, it is cheap to satisfy, and it applies to almost everyone using AI at work. It is <a href="https://artificialintelligenceact.eu/article/4/">Article 4</a>, and it is about AI literacy.</p>
<h2>What it actually says</h2>
<p>The wording is broad on purpose. Providers and deployers of AI systems have to take measures to ensure, to their best extent, a sufficient level of AI literacy among their staff and other people who operate and use AI systems on their behalf. It tells you to account for people's technical knowledge, experience, education and training, and the context the systems are used in.</p>
<p>Notice what it does not do. It does not prescribe a curriculum, a certificate or a number of hours. That vagueness makes some compliance teams nervous, but it is also a gift: you get to decide what "sufficient" looks like for your people, as long as you can show you thought about it and did something real.</p>
<h2>Why the EU bothered</h2>
<p>The logic is that half the risk in AI comes not from the models but from the humans pointing them at the wrong problem. Someone pastes a client's personal data into a public chatbot. Someone trusts a confident answer that happens to be wrong. Someone automates a decision that should have had a person in the loop. None of that is fixed by better model weights. It is fixed by people who understand what the tool is and is not.</p>
<div data-callout><p>The cheapest compliance win in the whole AI Act is probably this one. A short, honest training module plus a record of who completed it covers most of what Article 4 asks for. Compare that to the machinery around high-risk systems and it is not close.</p></div>
<h2>What "sufficient" looks like in practice</h2>
<p>You do not need to turn your finance team into machine-learning engineers. Sufficient literacy for most staff means they understand a few things well enough to act on them. That the tool can be wrong and sounds equally confident either way. That some information should never go into it. That the output is a draft, not an oracle. That there are tasks where a human has to make the final call. And that the rules differ depending on whether the tool runs inside the company or sends data outside it.</p>
<p>That last point is where the shape of your deployment matters. If your assistant runs on your own hardware and nothing leaves the building, a big chunk of the "do not paste that here" anxiety simply goes away, and your training gets shorter and calmer. If staff are using a mix of external tools, the literacy programme has to spend real time on what is safe to send where. The <a href="https://digital-strategy.ec.europa.eu/en/policies/ai-office">EU AI Office</a> has signalled it will publish living guidance and examples on literacy, so keep an eye there.</p>
<p>One practical note. Because Article 4 is already in force and enforcement of the broader Act is ramping up, "we will get to training later" is not a great position to be caught in. It is the one duty you can close out this quarter.</p>
<p>As always, this is general information rather than legal advice. Talk to a qualified adviser about what an adequate literacy programme looks like for your sector and your risk profile.</p>]]></content:encoded>
    </item>
    <item>
      <title>Shadow AI is already inside your company</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/shadow-ai-is-already-inside-your-company/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/shadow-ai-is-already-inside-your-company/</guid>
      <pubDate>Thu, 05 Mar 2026 09:00:00 GMT</pubDate>
      <category>Blog</category>
      <category>AI in practice</category>
      <description>Your employees are pasting company data into AI tools you never approved. The 2025 breach numbers put a price on it, and the fix is not a stern email.</description>
      <content:encoded><![CDATA[<p>There is an AI strategy running inside your company right now, and you did not write it. Your employees did, one paste at a time. It has a name in the security world: shadow AI. And the numbers on it stopped being theoretical.</p>
<h2>What the data says</h2>
<p><a href="https://www.ibm.com/reports/data-breach">IBM's 2025 Cost of a Data Breach report</a> put some hard figures on the phenomenon. One in five organisations that suffered a breach had a shadow AI element involved. Those breaches cost meaningfully more than average, a premium measured in hundreds of thousands of dollars. And of the organisations that had an AI-related breach, the overwhelming majority had no proper access controls on the AI tools in question. The global average breach already runs to 4.44 million dollars. Shadow AI is a line item now, not a curiosity.</p>
<p>The mechanism is completely ordinary, which is why it is so hard to stop. A salesperson pastes the customer list into a chatbot to draft outreach. A developer drops in proprietary code to debug it. Someone in HR summarises a sensitive case. Each of them is just trying to do their job faster. And each of them has quietly moved company data onto servers you do not control, outside your logs, possibly retained, possibly used to train a model. No alarm went off, because the data left through a browser tab like any other web page.</p>
<h2>Why banning it does not work</h2>
<p>The instinct is to send a stern email and block the tools at the firewall. It does not work, for a simple reason: the tools are genuinely useful, and people will find a way. They will use their phones. They will use a personal account. They will route around whatever you put up, because the productivity is real and the deadline is real. Prohibition without a substitute just pushes the behaviour further into the dark, where you can see even less of it.</p>
<div data-callout><p>The uncomfortable truth about shadow AI. It exists because your people want a capable assistant and you have not given them a sanctioned one. The demand is not going away. The only question is whether it gets met by a tool you control or a tool you do not.</p></div>
<h2>The fix is a better default</h2>
<p>The durable answer to shadow AI is not a stricter policy. It is a better default. Give people an assistant that is at least as capable as the public one they are tempted by, that runs inside your walls so the data stays put, and that is genuinely easy to reach. When the sanctioned tool is good and the data never leaves, the incentive to sneak off to a public chatbot mostly evaporates. People were not trying to leak data. They were trying to get help. Give them help that is safe and they will take it.</p>
<p>That is a large part of why we build AI to run on the customer's own hardware. It turns the shadow AI problem inside out. Instead of chasing every unsanctioned tool your staff might reach for, you provide one good tool where the sensitive processing happens in the building, and the reason to go elsewhere goes away. You cannot email your way out of shadow AI. You can build your way out of it.</p>]]></content:encoded>
    </item>
    <item>
      <title>DORA and AI: what financial firms have to prove about their models</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/dora-ai-models-financial-services/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/dora-ai-models-financial-services/</guid>
      <pubDate>Fri, 27 Feb 2026 09:00:00 GMT</pubDate>
      <category>EU regulations</category>
      <category>DORA</category>
      <description>DORA has applied since January 2025 and it worries openly about cloud concentration. What that means for banks and insurers running AI on third-party clouds.</description>
      <content:encoded><![CDATA[<p>Financial firms have lived under <a href="https://eur-lex.europa.eu/eli/reg/2022/2554/oj">DORA</a>, the Digital Operational Resilience Act, since 17 January 2025. Most of the coverage focused on incident reporting and resilience testing. The part that matters most for AI is quieter and more structural: DORA is openly worried that European finance has become dependent on a handful of cloud providers, and it has built machinery to do something about it.</p>
<h2>The concentration-risk problem, stated plainly</h2>
<p>DORA created an oversight regime for what it calls critical ICT third-party providers. In November 2025 the European Supervisory Authorities named the first batch, and the list read exactly as you would guess: AWS, Microsoft Azure, Google Cloud, Oracle, IBM and others, as the <a href="https://www.eiopa.europa.eu/digital-operational-resilience-act-dora_en">supervisors</a> confirmed. Regulators pointed out that a large majority of EU financial entities rely on at least two of the big three cloud platforms for critical functions.</p>
<p>Think about what that means as a systemic matter. If most of the continent's banks and insurers run their critical workloads on the same three American companies, then those three companies are single points of failure for the whole financial system. An outage, a compromise or a geopolitical rupture at one of them is not one firm's problem. It is everyone's problem at once. That is precisely the risk DORA was written to manage.</p>
<h2>What DORA asks firms to do</h2>
<p>DORA rests on several pillars, but the ones that touch AI directly are ICT risk management and third-party risk. In practice a financial entity has to keep a register of its ICT third parties, assess concentration risk, ensure it has exit strategies, and be able to keep operating if a provider fails. If your AI runs on a cloud model, that provider goes in the register, and you have to be able to answer the awkward question: what happens to this capability if this provider is unavailable, and can we actually move?</p>
<div data-callout><p>The uncomfortable maths for a lot of firms. Their AI, their productivity suite and their core infrastructure often sit with the same one or two hyperscalers. That is not diversification. It is concentration wearing a diversified costume, and DORA is designed to make supervisors ask about it.</p></div>
<h2>Where on-premise changes the answer</h2>
<p>Running an AI system on your own hardware does not make DORA go away. You still have resilience obligations, you still test, you still report incidents. But it changes the third-party risk picture in a specific, favourable way. A model that runs inside your own data centre is not another critical dependency on an externally overseen hyperscaler. It shortens the chain. Your exit strategy for that workload is trivial, because there is nothing to exit. And the concentration-risk question a supervisor might ask about your AI has a clean answer: this one does not add to our exposure, because we own it.</p>
<p>For the ICT third-party register, fewer external suppliers means less to document, monitor and stress-test. For concentration risk, keeping a critical workload in-house is the opposite of piling more load onto a designated critical provider. None of this is a silver bullet, and a badly run on-premise system can be less resilient than a well-run cloud one. But the structural point holds: the shorter your supply chain, the easier DORA's third-party questions are to answer.</p>
<p>This briefing is general information and not legal or regulatory advice. DORA is detailed and enforced by sector supervisors, so work through your specific obligations with your compliance function and, where relevant, Finansinspektionen guidance.</p>]]></content:encoded>
    </item>
    <item>
      <title>What we mean when we say your data never leaves the building</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/what-your-data-never-leaves-the-building-means/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/what-your-data-never-leaves-the-building-means/</guid>
      <pubDate>Tue, 24 Feb 2026 09:00:00 GMT</pubDate>
      <category>Blog</category>
      <category>Product</category>
      <description>It is an easy phrase to put on a website. Here is the concrete version: what actually happens to a prompt, a document and an answer when the AI runs locally.</description>
      <content:encoded><![CDATA[<p>"Your data never leaves the building" is the kind of line that sounds great and means nothing until someone explains it. Marketing pages love it. So let me do the unglamorous thing and describe what it actually means, step by step, because the difference from cloud AI is concrete and it matters.</p>
<h2>Follow a single question</h2>
<p>Imagine an employee asks the assistant to summarise a contract. With a cloud tool, here is the journey that question takes. The contract and the question leave their laptop, travel across the public internet, arrive at a provider's data centre (often in another country), get processed on machines nobody at your company will ever see, and the answer travels back. Somewhere in that trip, your contract sat on someone else's computer. Maybe it was retained for a while. Maybe it was used to improve a model. You are trusting a contract and a policy that the contract stayed private.</p>
<p>Now the local version. The employee asks the same question. The contract and the question go to a server sitting in your own building, on your own network. The model runs there. The answer comes back. Nothing crossed the internet. Nothing landed on a machine you do not own. There was no third party in the loop to trust, because there was no third party. That is the whole difference, and it is not subtle.</p>
<h2>What "the building" really refers to</h2>
<p>The phrase is a bit literal. What we mean is your trust boundary: the network and hardware you control. For most customers that is a server in their office or their own data centre. It can be genuinely air-gapped, with no internet connection at all, for the most sensitive work. Or it can be connected for convenience but configured so that the AI workload itself never sends data out. The point is that you decide where the edge is, and the sensitive processing happens inside it.</p>
<div data-callout><p>A quick way to test any "private AI" claim, including ours. Ask one question: when I submit a prompt, does the text of that prompt ever travel to a computer the vendor controls? If the answer is yes, in any form, for any reason, then the data does leave the building, whatever the brochure says. If the answer is no, ask them to show you how they would prove it.</p></div>
<h2>Why this is more than a feeling</h2>
<p>Keeping data local is not just reassuring. It removes real, specific problems. There is no international data transfer to assess against the shifting rules that followed the <a href="https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:62018CJ0311">Schrems II ruling</a>. There is no foreign statute that can reach a server which never held your data. There is no provider that could log your prompts or change its retention policy next quarter. When a regulator, an auditor or your own security team asks the hard question, where does the data go, you get to give the shortest possible answer. It stays here.</p>
<p>We seal the network perimeter on boot, keep an immutable log of what the system did, and ship the tooling to demonstrate that nothing left. That is the machinery behind the slogan. But the slogan itself is simple, and it is the reason a lot of our customers picked local AI in the first place. Their most valuable data is exactly the data they cannot send away. So we built the version where they do not have to.</p>]]></content:encoded>
    </item>
    <item>
      <title>The EU Data Act in plain words: what it changes for cloud and AI</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/eu-data-act-plain-words/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/eu-regulations/eu-data-act-plain-words/</guid>
      <pubDate>Wed, 18 Feb 2026 09:00:00 GMT</pubDate>
      <category>EU regulations</category>
      <category>EU Data Act</category>
      <description>The EU Data Act has applied since September 2025. Here is what it actually changes for cloud switching, egress fees and foreign access to your data.</description>
      <content:encoded><![CDATA[<p>Most people who run IT in Europe have heard of the GDPR and, by now, the AI Act. Far fewer have read the <a href="https://eur-lex.europa.eu/eli/reg/2023/2854/oj/eng">EU Data Act</a>, which is a shame, because it quietly rewires two things that matter a lot when you buy AI: how easily you can leave a cloud provider, and what happens to your data when a foreign government comes asking.</p>
<p>The Data Act, formally Regulation (EU) 2023/2854, has applied across the Union since 12 September 2025. The <a href="https://digital-strategy.ec.europa.eu/en/policies/data-act">European Commission's own summary</a> frames it as a law about "fairness in the data economy." That is true but a bit abstract. Let me pull out the parts that touch anyone running software.</p>
<h2>Leaving a cloud provider is supposed to get easier, and cheaper</h2>
<p>If you have ever tried to move a few terabytes out of a hyperscaler, you know the drill. The data is easy to put in and strangely expensive to take out. Those egress fees were never really about the cost of bandwidth. They were a fence.</p>
<p>Chapter VI of the Data Act (Articles 23 to 31) goes after that fence directly. Providers have to remove the commercial, technical and contractual obstacles that stop you switching to a different provider, moving to your own on-premises infrastructure, or running several providers at once. And Article 29 sets a clock on the fees themselves: switching charges are being reduced during a transition period and must be gone entirely from 12 January 2027, as legal analyses from firms like <a href="https://www.gtlaw.com/en/insights/2025/9/cloud-switching-under-the-eu-data-act">Greenberg Traurig</a> spell out. After that date, moving your data out is supposed to cost you nothing.</p>
<p>Read that again with a procurement hat on. The EU has decided that lock-in through exit fees is a problem worth legislating away. If a regulator thinks the switching cost is the issue, the cleanest answer is the arrangement where there is no provider to switch away from in the first place.</p>
<h2>The part about foreign governments</h2>
<p>Article 32 is the one that does not get enough attention. It requires providers of cloud and edge services to put in place technical, legal and organisational measures to prevent unlawful access by third-country governments to non-personal data held in the EU, where that access would clash with EU or member-state law.</p>
<p>Note the careful wording: non-personal data. Personal data transfers are still governed by Chapter V of the <a href="https://eur-lex.europa.eu/eli/reg/2016/679/oj">GDPR</a>, which is a separate conversation. But the fact that the EU felt it needed a dedicated article about foreign state access to ordinary business data tells you how real the concern has become. The worry has a name in most people's minds, the US CLOUD Act, and the Data Act is the Union writing a partial answer to it into law.</p>
<div data-callout><p>A blunt way to put it. The Data Act spends a lot of ink helping you escape a cloud provider and shielding your data from foreign legal reach. Both problems disappear if the data sits on a machine in your own building. You cannot be locked into an infrastructure you own, and no foreign court order reaches a server that never touched a hyperscaler.</p></div>
<h2>What this means if you are adding AI</h2>
<p>Here is where it gets concrete for anyone rolling out an assistant or a document-search tool. Every prompt, every uploaded file, every retrieved passage is data leaving your control the moment it hits a cloud model's API. The Data Act does not ban that. It just keeps chipping away at the assumption that your data naturally lives in someone else's data centre.</p>
<p>If you run the model locally, the whole question of switching costs, egress fees and third-country access stops applying to your AI workload. Not because you found a clever contractual workaround, but because the data never became someone else's problem to begin with.</p>
<p>None of this is legal advice, and the Data Act interacts with the GDPR, the <a href="https://digital-strategy.ec.europa.eu/en/policies/data-governance-act">Data Governance Act</a> and sector rules in ways your counsel should map for your situation. But the direction of travel is not subtle. Europe is legislating for a world where you can pick up your data and walk. Running AI on your own hardware is simply the shortest version of that walk.</p>]]></content:encoded>
    </item>
    <item>
      <title>Estimating your real LLM token bill: a working methodology</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/research/estimating-enterprise-llm-token-bill/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/research/estimating-enterprise-llm-token-bill/</guid>
      <pubDate>Wed, 11 Feb 2026 09:00:00 GMT</pubDate>
      <category>Research &amp; white papers</category>
      <category>Economics</category>
      <description>Most teams underestimate what cloud AI will cost at scale because they price the demo, not the workload. A method for estimating the bill before you sign.</description>
      <content:encoded><![CDATA[<p>The most common way companies get their AI budget wrong is charming in its optimism. Someone runs a few queries in a demo, sees a cost of a fraction of a cent, multiplies loosely, and concludes AI is basically free. Then the real usage arrives and the invoice does not match the story. This paper is a method for estimating the bill honestly, before you commit to paying it every month.</p>
<h2>Start with tokens, not dollars</h2>
<p>Everything in per-token pricing rests on one conversion, and it is worth internalising. As a rough rule, English text runs about 1.3 tokens per word, so a 750-word page is roughly a thousand tokens, per <a href="https://help.openai.com/en/articles/4936856-what-are-tokens-and-how-to-count-them">OpenAI's own guidance</a>. Treat that as approximate, not exact. Code and technical writing run higher. Other languages, including Swedish, inflate the count noticeably because the tokenizer was optimised for English. And newer models sometimes use tokenizers that produce meaningfully more tokens for the same text. So build in headroom rather than pretending the conversion is a constant.</p>
<h2>The three multipliers people forget</h2>
<p>A demo query is cheap. A real deployment is expensive for reasons the demo never shows you.</p>
<ul>
<li>Output tokens cost several times more than input tokens across every major provider, and useful assistants generate a lot of output. Weight your estimate accordingly.</li>
<li>Retrieval over your own documents means every question quietly drags a large slab of context through the model. A one-line question can carry ten thousand tokens of retrieved material behind it, and you pay for all of them, every time.</li>
<li>Agentic workflows, where the model calls itself in a loop to finish a task, can multiply token use by ten or more against a single-shot query. This is where the scary numbers come from, and it is exactly the direction serious deployments are heading.</li>
</ul>
<p>Any estimate that ignores these three is pricing a toy and budgeting for a tool.</p>
<div data-callout><p>A quick sanity anchor. Illustrative studies of heavy assistant use put a busy knowledge worker somewhere around fifty thousand input tokens and fifteen thousand output tokens a day. Over a working month that is roughly a million input and a few hundred thousand output tokens per person. Cheap on a small team. Run the same maths across a few hundred people using it hard, all day, with retrieval and agents, and the annual figure stops being a rounding error.</p></div>
<h2>A method you can defend</h2>
<p>Here is a way to produce a number you will not be embarrassed by later. Pick two or three real roles, not averages, and estimate their honest daily token use including retrieval and any agentic workflows. Convert to a monthly per-role figure. Multiply by how many people in each role will actually use the tool, not the whole headcount and not just the pilot group. Apply the provider's real input and output rates separately, because the blended number hides the expensive half. Add a growth factor, because usage of a tool people like goes up, not down. Then annualise. The result is usually two to five times the back-of-napkin figure the demo suggested, and it is the number worth taking to a decision.</p>
<h2>Why this exercise favours owning the hardware, sometimes</h2>
<p>Run this method and one of two things happens. Either the annual cloud number is modest, in which case the cloud is the right answer and you should stop reading vendor white papers like this one. Or it is large, and it recurs every year forever, and it grows. In that second case, the comparison against owned hardware, which is a fixed cost that does not scale with every extra query or every new hire, suddenly looks different. We walk through that comparison in our companion piece on <a href="/research/on-premise-ai-tco-versus-cloud/">the total cost of ownership of on-premise AI</a>. The point of this paper is narrower: get the real cloud number first, because most people never do, and it is the number that decides everything else.</p>
<p>Figures here are illustrative and drawn from published rules of thumb and vendor rates current in 2026. Your actual usage is the only estimate that matters, which is exactly why it is worth measuring before you sign.</p>]]></content:encoded>
    </item>
    <item>
      <title>The hidden cost of per-token pricing</title>
      <link>https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/hidden-cost-of-per-token-pricing/</link>
      <guid isPermaLink="true">https://delightful-grass-0e0338903.7.azurestaticapps.net/blog/hidden-cost-of-per-token-pricing/</guid>
      <pubDate>Wed, 28 Jan 2026 09:00:00 GMT</pubDate>
      <category>Blog</category>
      <category>Economics</category>
      <description>Per-token AI looks almost free in the demo. The trouble is what happens when a whole company uses it every day, for years, and the meter never stops.</description>
      <content:encoded><![CDATA[<p>The first time you use a cloud AI API, the price is intoxicating. You ask it something useful and it costs a fraction of a penny. Your brain does the obvious multiplication, decides AI is basically free, and moves on. That moment of arithmetic is where a lot of expensive decisions get made.</p>
<p>Here is what the demo does not show you.</p>
<h2>The meter never stops</h2>
<p>Per-token pricing is a meter. It runs on every single interaction, forever. That is fine when three people use it now and then. It is a very different thing when four hundred people use it all day, every working day, for the next five years, and the meter is still running the whole time. You are not buying a tool. You are renting a habit, and the rent is due every month whether the tool got more valuable or not.</p>
<p>Owned things have a different shape. You pay once, and then usage is close to free. A meter has the opposite shape. Usage is the cost, and the more your people rely on AI, the more you pay for their reliance. Success makes the bill go up. Think about how strange that incentive is.</p>
<h2>The half of the bill nobody quotes</h2>
<p>There is a detail in the pricing that quietly doubles or triples real costs. Output tokens, the text the model writes back, cost several times more than input tokens across every major provider. Look at the published rates from <a href="https://platform.claude.com/docs/en/about-claude/pricing">Anthropic</a> or <a href="https://ai.google.dev/gemini-api/docs/pricing">Google</a> and you will see the same pattern: output is where the money is. A good assistant is one that writes a lot back. So the better the tool does its job, the more of the expensive kind of token it generates. The pricing punishes exactly the behaviour you wanted.</p>
<p>Then there is retrieval. The moment your assistant answers questions using your own documents, every question drags a big slab of context through the model, and you pay for all of it. A one-line question can carry ten thousand tokens of retrieved material behind it. And agentic workflows, where the model works in loops, can multiply the count tenfold. None of this shows up when you are poking at a demo with single questions.</p>
<h2>Do the annual maths, once</h2>
<p>I am not going to tell you the cloud is always the wrong choice, because it is not. For light, occasional use, the cheap model tiers are genuinely hard to beat, and you should use them. The problem is not the cloud. The problem is signing up for a per-token meter without ever working out what it costs at full tilt.</p>
<p>So do it once. Take a real heavy user, estimate their honest daily token use including retrieval and any agents, multiply by the people who will actually use the tool hard, apply the real input and output rates separately, add growth because usage of a good tool goes up, and annualise. We walk through the method in more detail in our note on <a href="/research/estimating-enterprise-llm-token-bill/">estimating your real token bill</a>. Most people never run that number. It is usually two to five times the figure the demo implied.</p>
<p>Once it is in front of you, one question does the rest of the work. Would you rather pay that number every year, growing, forever? Or own the thing that produces the same answers, on hardware in your building, at a fixed cost that does not care how many questions your people ask? For a lot of companies, seeing the annual figure written down is the entire decision.</p>]]></content:encoded>
    </item>
  </channel>
</rss>
