AI infrastructure in banking: where the compute sits

A bank that wants its own models faces three decisions before any model is chosen, and every one of them is an infrastructure decision. Does the bank need to train or only to run a model that someone else trained. Does the compute sit in its own data center, in a sovereign cloud, or at a hyperscaler. And are the model weights something the bank holds or something it calls over an API. The answers determine the cost, the supervisory file and what an auditor can be shown two years later.

Finance Loop covers what banks do with AI at AI in banking in Germany and the European picture at AI in finance in Europe. This is the layer underneath those use cases.

GPU accelerator racks connect to visible liquid-cooling pipes in a bank data center.

Inference against training, and which one a bank needs

Almost no bank needs to train a foundation model, and the ones that believe they do usually need something else. Training from scratch means thousands of accelerators for weeks, a data pipeline nobody in the house has built, and a result that trails the open releases. Fine-tuning an existing model on the bank's own documents is a different scale: tens of accelerators for hours, repeatable, and the output is a model the bank controls.

Inference is the workload that actually recurs. It runs every time a credit analyst asks a question of a document set or a service case gets summarized, and it has to meet a latency the business set, not one the vendor quotes. Real-time fraud scoring is where the two part company: a decision that has to land inside a payment authorization cannot wait for a round trip to an external API, and that single requirement pushes the workload onto infrastructure the bank controls. Retrieval over the bank's own data makes most answers better, where another training run would not, because the model was never the part that lacked the bank's knowledge.

Own data center, private cloud, sovereign cloud, hyperscaler

Four places, four different sets of problems. The bank's own data center gives the strongest control story: the model, the inference layer and the data pipeline sit in a facility the bank owns or leases, and customer data never crosses the perimeter. The price is accelerator procurement, power and cooling for racks that draw far more than the servers they replaced, and a hardware refresh cycle the bank now owns.

A sovereign cloud puts the workload with an operator bound by jurisdiction: contractual and in places legal commitments that data, encryption keys and operational control stay inside a named legal space, run by an entity under that law. Finance Loop's page on the subject is sovereign cloud for banks. Between the two sits the private cloud or virtual private cloud: dedicated tenancy inside a large provider, with strong logical isolation and provisioning in weeks instead of quarters. Sphere's comparison of the three models puts the timelines side by side, with six to twelve months or more for an own build against weeks for dedicated tenancy, and it notes that the sovereign option carries a compliance premium over the plain private cloud. A hyperscaler gives the capacity and the model choice immediately, and leaves the bank arguing about which entity holds the keys and which staff outside the EU can reach the control plane.

The European answer to that argument is under construction. Lloyds Banking Group and NatWest Group have joined a coalition around a UK sovereign frontier model so that banks can run the model inside their own infrastructure instead of sending customer data to a non-domestic API. The same question drives the sovereign AI discussion in Germany, where it meets the Gaia-X work and the EU Cloud Code of Conduct.

Open weights against a closed API, and what each does to the audit trail

An open-weight model is a file the bank can hold. That makes three things possible that a closed API does not: the exact model version can be frozen for as long as a supervisor might ask about a decision, the model can run with no outbound network connection, and the bank can evidence that the weights did not change between the decision and the review. A closed model behind an API can be updated by its provider, and a decision made against last quarter's version cannot be reproduced once that version is retired.

Closed models give capability the open releases often trail, plus a vendor who carries the operational burden. The practical pattern in regulated houses is both: a closed model for drafting and research where no decision hangs on it, and a self-hosted open-weight model where a decision does. The deciding question is not which model scores better on a benchmark, but which decisions have to be reconstructable, and for how long. Finance Loop treats the model layer at large language models in finance and generative AI in finance.

Data residency for prompts and outputs

A prompt that carries a customer's name is a transfer of personal data, and it is a transfer each time, not once at the start of the project. GDPR applies to the prompt, to any logs the provider keeps of it, to the output and to the retrieval index built from customer documents. The questions that settle it are concrete: where the inference endpoint physically runs, how long the provider retains prompts and for what purpose, whether the data is used to improve the provider's model, and which staff of the provider can read a prompt in the course of support.

The retrieval index is the part firms forget. A vector database built from customer correspondence holds the same personal data as the correspondence, is subject to the same deletion duties, and has to be rebuilt when a record is erased. A deletion request that removes the source document and leaves its embedding is not a deletion.

DORA applied to a model provider

A model provider is an ICT third-party service provider, and DORA treats it as one. That means the contract carries the mandatory terms on access, audit and exit, the provider appears in the register of information the firm maintains, and the firm has to be able to say what happens to the function if the provider stops serving it. For a model this is harder than for a database, because an exit plan has to name the substitute model and say how the prompts, the retrieval layer and the evaluation set move to it.

Concentration is the second question DORA asks. A bank whose fraud scoring, customer service and credit summarization all run against one provider has one dependency, however many applications it counts. Finance Loop's page on the German supervision of this is DORA in Germany, and the cloud background sits at cloud computing in banking.

What the EU AI Act expects on file

Credit scoring of natural persons is a high-risk use under the EU AI Act, and life and health insurance pricing is as well. For those systems the documentation has to exist before deployment: what the system does, which data trained or configured it, what the known limitations are, how human oversight works in the process, and what the system scored for accuracy and for resilience to disturbed input on the test set. Logs of the system's operation have to be kept.

The infrastructure decision shows up in that file directly. A self-hosted model with a pinned version can produce the artifacts on request; a closed API whose version the provider rotates needs a contractual commitment on version retention and log access, obtained before deployment and not during the first supervisory question. Finance Loop covers the regime at the EU AI Act in finance.

Capex, opex and the hybrid answer most banks land on

The cost shapes differ more than the totals. An own build is front-loaded capital: accelerators bought before the first user, data center capacity, physical security and staff who know how to run the hardware. Cloud tenancy is operating expense that scales with use, which suits a workload nobody has sized yet and punishes one that runs flat at high volume for years. The sovereign variant adds a premium over plain dedicated tenancy for the jurisdictional guarantees.

What this produces in practice is a portfolio and not a single choice. The deployment decision belongs to the data class and not to the institution: the workloads that touch the most sensitive customer data go to the strictest environment, and everything else goes where provisioning is fast. Sphere's framework makes the same point from the cost side, warning against paying for control the bank does not need everywhere. A bank that picked one environment for all of its AI either overpaid for the easy workloads or placed the hard ones wrongly.

The decision order that holds up is regulatory first and commercial second. A mandate that data never leaves the building decides for an own data center. A residency requirement decides for a sovereign environment. A requirement for isolation and auditability without residency is usually satisfied by dedicated tenancy. Only when none of the three binds does the question become one of price and speed.

What is AI infrastructure?

AI infrastructure is the compute, storage, network and software layer that model training and model inference run on: accelerators and the hosts around them, high-bandwidth storage for the data and the model files, a network fast enough that the accelerators are not waiting, and the orchestration, serving and monitoring software above all of it. In a bank it also includes the identity, logging and entitlement systems that make the use auditable.

Can a bank run a large language model on its own servers?

Yes, and many do for the workloads where the data cannot leave. An open-weight model in the usual sizes runs on a single server with a few accelerators, which is an ordinary procurement and no data center project. What grows is not the hardware but the surrounding work: version control for weights, an evaluation set that is rerun on every change, capacity planning for concurrent users, and a monitoring layer that catches a degraded answer before a user acts on it.

What does sovereign AI mean in finance?

Sovereign AI in finance means the models, the data and the compute stay under a jurisdiction's law and under operators subject to it, so that no foreign access regime reaches customer data and no foreign decision removes the capability. For a European bank it covers three separate things: where the data sits, who holds the encryption keys and who can reach the control plane. A deployment can satisfy one of the three and fail the other two, which is why the term needs unpacking in every contract.

How much GPU capacity does a bank need for its own model?

For inference on an open-weight model of the usual sizes, one server with a small number of accelerators serves a department, and the sizing question is concurrency and not model size: how many users expect an answer in the same second, and how long the answers are. For fine-tuning on the bank's own documents, a handful of accelerators for hours per run is the realistic figure, and the run repeats whenever the document base changes. Training a foundation model from scratch is a different order of magnitude and is not what a bank should be planning. The capacity that gets forgotten is the second environment: a production deployment needs a test deployment with the same model version, or there is nowhere to validate a change.

What is the difference between private cloud and sovereign cloud?

A private cloud gives dedicated tenancy and logical isolation inside a provider's platform, so the bank's workload does not share compute with other customers. A sovereign cloud adds commitments about jurisdiction: that the data, the encryption keys and the operational control stay in a named legal space, with an operator subject to that law. The difference shows up when a foreign authority asks the provider for data. Logical isolation does not answer that request; a jurisdictional commitment is written to.

AI infrastructure in banking and Finance Loop

Finance Loop is the meeting place in Frankfurt for the Digital Infrastructure & Sovereignty track, where the people who sign the cloud contracts, the architects who size the accelerators and the risk managers who answer for the dependency sit in the same room. Finance Loop members hear what an infrastructure decision cost somebody else before they make their own. Finance Loop keeps the discussion on the operating reality and not on vendor roadmaps.

Finance Loop is a professional network and has the goal of driving the adoption of emerging technologies in finance, such as AI, tokenization, stablecoins, and DeFi. Finance Loop helps its members build skills and personal networks in these fields: Investment & Digital Assets, Payments & Digital Money, Digital Infrastructure & Sovereignty, and Risk & Compliance.

Let's stay in touch

4,000+ members in finance and tech. Become a Network Member for free.

Get updates for free!

Exclusive event invitations, member perks and news from the network. Unsubscribe at any time.

By submitting you agree to the terms.