Contact Us

Deployed On-Premise GenAI Inside a Bank in 3 Weeks — Starting with Document Intelligence

Itexus deployed an on-premise LLM inside a security-constrained bank in 3 weeks — automating document classification, PDF bundle splitting, and field extraction, while defining a scalable AI architecture for future use cases.

Document Intelligence dashboard of the on-premise GenAI solution shown on a laptop
Document Intelligence dashboard of the on-premise GenAI solution shown on a laptop

Case Summary for Financial Services Leaders

A bank in a high-security regulated environment needed to adopt GenAI without using public SaaS AI services due to data confidentiality constraints. Building a full in-house AI platform from scratch was too slow and expensive for a first step. Itexus proposed a practical entry point: deploy an open-weight LLM on-premise inside the bank's own infrastructure, validate the model on a real document workflow, and define a scalable AI architecture for future automation.

Technologies: Python, PyTorch, HuggingFace Transformers, FastAPI, Docker.

Results: production-ready MVP in 3 weeks, 20 document types supported, 95% classification accuracy, ~2 seconds per page.

IndustryBanking
SegmentConsumer Bank
GeographyNorth America
Deployment ModelOn-Premise (air-gapped / private infrastructure)
Duration3 weeks (MVP scope)
Team Size2 (1 AI-Native Product Engineer + 1 Project Manager)
Delivery ModelAI-first SDLC
ServicesAI-first product engineering, Banking software development
TechnologiesPython, PyTorch, HuggingFace Transformers, FastAPI, Docker, PostgreSQL
CompliancePCI DSS, SOC 2

About the Client

A regulated banking institution operating in a high-security environment processes sensitive customer documents across multiple internal operational workflows daily. The bank's security and compliance policies prohibit the use of externally hosted AI services — including public cloud LLM APIs — for any workloads involving customer data.

The bank required an AI adoption path that would fit within its existing controlled infrastructure, comply with strict data residency and confidentiality requirements, and be realistic as a first step rather than a multi-year platform investment.

The Challenge

The challenge was not technical in the narrow sense. The real question was: how does a security-constrained bank take a first step into GenAI adoption — without compromising its data perimeter and without launching a large, expensive platform program?

Business Challenge

  • Public SaaS AI tools — including major cloud LLM APIs — were off-limits for workloads involving sensitive customer documents.

  • Commissioning a full proprietary AI platform from scratch was too slow and too capital-intensive for a first step.

  • Existing document workflows were still operationally painful: manual classification, no bundle splitting, no field extraction.

  • The bank needed a pilot that proved GenAI could be run safely inside its own perimeter — not as an isolated experiment, but as a realistic operating model.

Technical Challenge

  • Deploying a capable open-weight LLM on-premise inside the bank's controlled infrastructure without relying on external inference APIs.

  • Building page-level document classification capable of handling multi-page PDFs and bundled files containing 5–10 consecutive documents.

  • Implementing document boundary detection inside PDF bundles — a non-trivial structural problem distinct from simple classification.

  • Designing structured field extraction with JSON output compatible with the bank's downstream processing systems.

  • Architecting the MVP as a foundation for future AI use cases — not a standalone utility — including an API Gateway, Document Processing Layer, and AI/ML Services Layer.

Solution

Itexus delivered the project through an AI-first SDLC model — a lean, AI-native delivery approach that enabled a 2-person team to move from concept to production-ready MVP in 3 weeks. The solution was deployed entirely on-premise inside the bank's protected infrastructure.

Document Intelligence MVP

  • Ingestion of multi-page PDF files from internal bank systems

  • Handling of document bundles containing 5–10 consecutive documents within a single PDF

  • Page-level document classification — each page is classified individually before boundary detection

  • Document boundary detection inside bundled PDFs — identifying where one document ends and the next begins

  • Structured field extraction from classified documents

  • JSON output formatted for integration with the bank's existing downstream systems

  • Scope expanded from 10 to 20 document types at the MVP stage

On-Premise Deployment & Security Controls

  • Sensitive customer documents remained within the bank's controlled environment at all times

  • The pilot operated as a realistic model of future AI adoption, not an isolated proof-of-concept disconnected from real infrastructure

  • Security controls included authenticated access, role-based permissions, auditability of processing events, and automatic cleanup of temporary processing artifacts

Target Architecture for Future Scale

  • API Gateway: a unified entry point for all AI service requests across the bank's internal systems

  • Document Processing Layer: orchestration of ingestion, classification, boundary detection, and extraction pipelines

  • AI/ML Services Layer: model serving infrastructure capable of hosting multiple models and supporting future automation scenarios

Use Cases Covered

The MVP addressed the following document automation use cases within the bank's internal workflows:

Use CaseDescriptionOutput
Document ClassificationPage-level identification of document type across 20 classesDocument type label per page
PDF Bundle SplittingDetection of document boundaries inside multi-document PDFsSeparated individual document records
Structured Field ExtractionExtraction of key fields from classified customer documentsStructured JSON
Downstream IntegrationJSON output formatted for bank's existing internal systemsAPI-ready structured data
Audit LoggingProcessing events logged with role-based access controlsTraceable processing history

Problems Solved

  • Document processing time reduced by 40% with data extraction accuracy up to 95%.

  • Enabled safe GenAI adoption inside a regulated bank. The bank now has a validated, working model for running LLM-based automation inside its own infrastructure — without relying on public AI services or building a full platform program.

  • Expanded document type coverage. The system went from 10 to 20 supported document classes, with an architecture designed to extend further without fundamental rework.

  • Delivered a reusable AI infrastructure foundation. The MVP is not a standalone tool. The target architecture defined alongside it provides a platform for future AI use cases across the bank's operations.

Solution Architecture

The system is organized across three functional layers:

LayerComponentsFunction
API GatewayREST API, Auth, Rate Limiting, Request RoutingUnified entry point for all AI service requests
Document Processing LayerPDF Ingestion, Page Classifier, Bundle Splitter, Field ExtractorEnd-to-end document workflow orchestration
AI/ML Services LayerLLM Inference Engine, Model Registry, Artifact StorageModel serving, reusable for future AI use cases
Security & ComplianceIAM, Audit Logging, Artifact Cleanup, Role-based AccessControls required in a regulated banking environment
On-Premise GenAI infrastructure for a Bank
On-premise deployment - all layers run inside bank's controlled infrastructure

Integration Landscape

Integration points documented in the project:

  • Internal document management systems (source of PDF input)

  • Bank's downstream processing systems (consumer of JSON output)

  • Internal IAM and audit infrastructure

  • Open-weight LLM model (deployed on-premise, not via external API)

Technology Stack

CategoryTechnologies
AI / MLHuggingFace Transformers, PyTorch, open-weight LLM (on-premise)
BackendPython, FastAPI
Document ProcessingPDF parsing libraries, custom classification and extraction pipeline
InfrastructureDocker, on-premise deployment
DatabasePostgreSQL
SecurityIAM, role-based access control, audit logging

Compliance & Security

Security architecture was a first-class design constraint, not an afterthought. The solution was built specifically for a regulated banking environment with strict data residency requirements.

Security Measures Implemented

  • On-premise LLM deployment — no data leaves the bank's infrastructure

  • Authenticated access with role-based permission controls

  • Auditability — all document processing events are logged

  • Automatic cleanup of temporary processing artifacts

  • No external API calls during document processing

  • Controlled model deployment — model runs on bank's own hardware

Compliance Considerations

  • Data residency — customer documents remain within the bank's controlled perimeter

  • Internal security policy compliance — no public cloud or SaaS AI dependencies

  • Auditability requirements for regulated financial institutions

  • Architecture designed for future alignment with broader banking compliance frameworks (e.g. GDPR, sector-specific AI governance)

Delivery Approach

Delivery Model: AI-first SDLC

Team Composition

RoleScope
AI-Native Product EngineerArchitecture, model selection, pipeline implementation, on-premise deployment, security controls
Project ManagerClient communication, scope management, delivery coordination

Governance

  • CI quality gates on all model outputs and pipeline components

  • Clear AI usage rules defined in repository

  • Restricted access to high-risk areas (production data, security controls)

  • Traceability and audit logs for all processing events

  • Scoped model context — no unnecessary data exposure to the model

  • Reference code patterns to ensure architectural consistency

  • Domain-based access controls aligned with bank's security requirements

  • No direct production changes — all changes via controlled deployment pipeline

  • Risk-based controls applied to security-sensitive components

Results

The bank received more than a working MVP. It received a validated, repeatable model for running GenAI inside its own environment — and a defined architectural path to extend AI automation to other workflows.

3 weeks
concept to production-ready MVP
95%
classification accuracy
~2 sec
per page processing speed

How This Expertise Can Help Your Team

The capabilities demonstrated in this project map directly to several recurring challenges in financial services AI adoption:

AI-first DevelopmentThe delivery model used in this project. Lean teams, AI-native tooling, fast time-to-value on complex financial systems.
AI Software DevelopmentBuilding AI-powered features inside banking products — classification, extraction, automation, intelligent workflows.
Banking Software DevelopmentEngineering for regulated banking environments: security, compliance, integration with core banking systems.

FAQ

What type of organizations is this solution most relevant for?

Banks, insurance companies, and other regulated financial institutions that process high volumes of customer documents and cannot use public cloud AI services due to data security or regulatory constraints. Also relevant for any financial institution looking to validate GenAI adoption without committing to a large platform program upfront.

What does 'AI-first SDLC' mean in practice, and how is it different from standard AI development?

AI-first SDLC means AI tooling is integral to the engineering workflow itself — not just used for specific tasks like code completion. In this project, it enabled a 2-person team to deliver in 3 weeks what a traditional team would typically take 3–6 months to scope, design, and build. The key difference is that the team structure, process, and toolchain are all designed around AI augmentation from day one.

Why was on-premise deployment important — couldn't this be done in a private cloud?

For this bank, deploying on private cloud infrastructure still meant relying on external service providers with access to the infrastructure layer. The requirement was for the model and all processed data to remain within the bank's own controlled environment — no external dependencies during inference. On-premise deployment satisfied this requirement; private cloud VPC configurations may or may not, depending on the bank's specific policies.

What were the biggest technical risks, and how were they managed?

The primary technical risk was model performance on the bank's specific document types, which differ from typical benchmark datasets. This was addressed by benchmarking the model against the bank's internal document set before committing to the approach, and by designing the pipeline to support model swapping without architectural changes. The second risk was document boundary detection in bundled PDFs — a problem that requires the model to reason about page sequences, not just individual pages. This was handled as a separate pipeline stage with dedicated validation.

How was the migration executed without disrupting operations?

The MVP was deployed as a new, parallel system — it did not replace existing document workflows during the pilot phase. This allowed the bank to evaluate the system on real documents against known-good results before committing to operational integration.

What happened after the MVP, and what is the next step?

The MVP delivered two things: a working production system for document intelligence, and a validated target architecture for future AI use cases. The next steps for the bank involve extending the document types covered and using the defined architecture to onboard additional automation scenarios — such as compliance document review or onboarding verification — onto the same AI infrastructure.

What lessons from this project apply to other banks considering GenAI adoption?

Three consistent patterns emerged: (1) Start with a real workflow that is currently painful, not a synthetic proof-of-concept. (2) Design for your security constraints from day one — retrofitting security onto a cloud-first architecture is more expensive than building on-premise from the start. (3) Use the first use case to define a reusable AI architecture — the pilot should be the first workload on a platform, not a one-off tool.

What factors most affect success in banking AI adoption projects?

Security architecture is the most commonly underestimated factor in regulated banking environments. Teams that treat it as a first-class constraint (not a compliance checkbox) reach production faster. The second critical factor is team composition: AI-native engineers with prior experience deploying models in regulated environments dramatically compress timelines compared to teams learning on the job.

Looking for a Fintech Engineering Partner?

We help financial institutions adopt AI safely, build fintech products that scale, and deliver engineering outcomes — not just code.