On-prem in under 5 minutes: Jina embedding models now available for on-prem deployment

All 28 Jina AI models, including rerankers, as ready-to-deploy Docker containers, with zero telemetry and no license server. Drop-in compatible with OpenAI, Cohere, Voyage AI and Elastic Inference Service APIs.

Get hands-on with Elasticsearch: Dive into our sample notebooks in the Elasticsearch Labs repo, start a free cloud trial, or try Elastic on your local machine now.

All 28 Jina AI embedding and reranking models now ship as fully offline Docker containers for on-prem deployment, including jina-embeddings-v5-omni and jina-reranker-v3. Download one, transfer it to an on-premises air-gapped or firewalled system, and local inference is running in under five minutes. The containers are completely self-contained and make no external connections. There’s no call to Hugging Face or any model registry. There’s also no license server or telemetry or logging endpoints. For regulated industries, data sovereignty requirements or environments where internet access is unreliable or simply unavailable, this removes the dependency on third-party AI services. Jina On-Prem supports Elastic Inference Service (EIS), OpenAI, Cohere, Voyage AI, and Gemini API schemas, so existing applications work without code changes.

The most powerful AI models run on remote cloud installations with access via a web API, meaning that you have to trust your AI service provider for security, service availability, and stable prices. You can’t easily align reasonable demands for reliability, privacy, manageable costs, and good data governance with increasingly powerful, sophisticated, and resource-intensive AI usage.

Government regulation, court rulings, and business considerations made in someone else’s interest have all recently resulted in restricting access to specific services. And even if you can switch to other services, AI models aren’t components that can just be swapped out whenever you want. Applications that use semantic embeddings depend on having access to the same models at query time as at data ingestion time. To lose access to your embedding model means your search system comes to a halt.

AI pricing models compound that risk. Recent financial disclosures from major AI vendors give customers good reason to be concerned about potential price hikes. Reliance on products with unpredictable costs adds more risk to capital-intensive AI investments that may not produce clear returns.

Jina On-Prem is Elastic’s answer to these challenges.

Who needs on-premises AI?

Local hosting and direct control over your AI models support a variety of technical demands, industry requirements, and business interests.

Local installation reduces what you pay your AI service providers, but it puts the cost of hardware and reliable access on your organization. Depending on your volume of use, it may simply be cheaper. But there are additional pressing reasons to consider running your own AI. If any of the issues described below concern your enterprise, consider a local AI solution like Jina On-Prem. This list is not exhaustive.

Use caseWhy on-premExample
Air-gapped / high-securityNo outbound data transmission; complete network isolationDefence, intelligence, classified research
Regulatory complianceData sovereignty; no cross-border transmission or third-party exposureHealthcare (Health Insurance Portability and Accountability Act [HIPAA]), finance, EU enterprises (General Data Protection Regulation [GDPR])
Latency-criticalZero network dependency; no tolerance for connection failuresRobotics, edge computing, vehicles, ships
Cost predictabilityFixed infrastructure cost vs. per-token pricing with uncertain future ratesHigh-volume continuous inference workloads
Liability reductionNo third-party data exposure; maintains legal privilege and duty of careLaw firms, government agencies

Why air-gapped and firewalled systems need on-prem AI

Air-gapped and firewalled systems cannot use external AI APIs. Jina On-Prem runs entirely within your infrastructure with no outbound connections.

For organizations managing especially sensitive data, security and privacy considerations are paramount. It does little good to invest in protecting your sensitive data if you promptly turn it over to some remote third party that may have insufficient security in place or might be subject to the demands of a foreign government.

Employees in organizations that handle sensitive data often receive some training in secure data handling, but this isn’t very effective when they all have web browsers that may be open to any page on the internet while they handle that data. Isolation is the most effective security measure available, either through air-gapping or very restrictive firewalls, but that makes it difficult to use external services of any kind.

On-prem AI for latency-sensitive and high-availability systems

Software as a service and cloud computing represent a compromise between the cost of offering highly accessible, reliable services on your own computers and outsourcing the problem to someone else. But they come with variable latency, outages, and a complete loss of control when things go wrong. AI services aren’t the exception. If your search system goes offline when you can’t access your embedding model, it may no longer look like a good compromise.

Furthermore, relying on external AI will always involve risks that you can’t easily foresee or manage. Internet access and network latency can degrade without notice, as a result of political events, bad weather, or ships dragging their anchors over underwater fiber-optic cables. Governments can, and recently have, used export bans to suddenly block access to AI models. AI service providers sometimes withdraw models to induce you to switch to newer ones. The flexibility and managed costs of external services have to be balanced against the risks of dependency.

On-prem AI for GDPR, HIPAA, and data sovereignty compliance

Organizations that collect personal data are subject to increasingly stringent regulations which often differ between jurisdictions and may have contradictory requirements. Notably, HIPAA rules place very strict data protections on American healthcare providers, and strong general data protection laws in Canada, the European Union, and many Asian jurisdictions require all enterprises that handle personal information to do so securely and to limit the transmission of that data to other parties or other jurisdictions. These rules can even impose obligations on foreign entities if they have any customers in those jurisdictions. Financial institutions are frequently subject to even stricter rules and bear the same direct liability for information security that they have to protect against other forms of criminal activity.

Regulatory compliance can be incompatible with third-party AI services, especially if using them involves cross-border data transmission.

Furthermore, recent events show that rules restricting the physical location of data stores may not be a reliable source of protection when international cloud operators are subject to pressure from foreign governments. Local laws may conflict between jurisdictions, requiring local data storage and processing and making third-party services impossible to use. In some cases, the only solution is to take all the parts of your processes in house, including your AI systems.

AI liability risks from third-party data transmission

Data protection laws and recognized duties of care toward sensitive data routinely have liability implications, sometimes very severe ones. You can be liable for third-party service providers’ handling of your data. While courts and legal procedures might provide some retrospective protections from insecure service providers, those remedies are not available nor generally effective against national security actors, law enforcement, or criminal hackers.

For governments, there have already been instances of cross-border cloud service providers releasing sensitive state information to foreign actors.

But even if you don’t worry about foreign governments or hackers, and if your external AI service providers are themselves secure, just the fact that they’re external can create liabilities.

For example, in most jurisdictions, lawyers’ communications with their clients enjoy special legal protections, and law offices have strict liabilities when recording or storing this information. In the United States, this “attorney-client privilege” is so famous, it’s central to movie and TV plots. But one of the ways that privilege can be lost is by communicating information with someone who is not privileged, and recent developments suggest that external AI service providers might qualify.

It’s possible, at least in the United States, that just using third-party AI services over an internet API, like embedding models that provide indexing services, might violate critical confidentiality rules. A law firm might be sued, disciplined, or disbarred just for using externally hosted software, even if no security breach occurs.

On-prem AI for offline, edge, and physically isolated systems

Computer systems aren’t just isolated for security reasons. For example, moving vehicles cannot rely on internet access for any essential functions. Ships and aircraft have very extensive onboard computer systems that have to function without internet connections and therefore cannot use external AI services. Offshore platforms, remote facilities in wilderness areas, computer services in the Arctic, Antarctic and on small islands without adequate physical connections to global networks are all examples of installations that benefit from locally hosting all the services they need. As AI’s role in enterprise computing grows, these limitations become more important to address.

Emerging applications of AI to physical systems (robotics and other spatially confined or external-world–focused use cases, like logistics management systems or even supermarket checkouts) may be connected to the global internet, but they have no tolerance for connection failures or spikes in latency. If they rely on an AI system to operate, that AI system needs to be as local and reliable as possible.

Who doesn’t need on-premises AI?

Remote software services and off-site AI do have benefits. Running AI models can require expensive, power-hungry processors with notoriously short lifespans. Access to high-quality hardware is particularly difficult right now due to market factors and external economic shocks. Under the circumstances, it may make sense to pay by the token to use an external API instead of supporting the steep capital costs of local AI.

External APIs make the most sense for intermittent users. If you use AI models primarily to batch process data for analysis, rather than running a search system that has to be online all the time, it makes little sense to invest in capital-intensive hardware and local installations.

Furthermore, when your data processing is already cloud-based, for example, an ecommerce website hosted in the cloud for reliability and accessibility reasons, using AI services located in the same cloud infrastructure may provide a better value for money than introducing your own licensed AI model deployment. You’re already dependent on your cloud service provider, so being dependent on its AI services doesn’t add much risk.

If your use case sounds like it fits that description, Jina AI models are available on EIS, AWS Marketplace, and the Google Cloud Platform specifically to meet your needs.

The table below summarizes the key factors. Your answer depends on your data, infrastructure and usage pattern.

FactorOn-prem favoredCloud API favored
Usage patternContinuous or high-volume inferenceIntermittent or batch processing
Data sensitivityRegulated, sovereign, or classifiedNo cross-border or third-party restrictions
Network environmentAir-gapped, firewalled, or unreliableStable, always-on internet
Existing infrastructureOwn or can procure GPU hardwareAlready cloud-hosted with colocated AI
Cost modelFixed hardware + license; predictable at scalePer-token; lower up-front, variable long-term
Latency toleranceNone (robotics, edge, real-time)Network variability is acceptable
Operational responsibilityYour team manages hardware and availabilityProvider manages hardware and updates; you manage integration

You have to consider the costs and benefits in light of your particular circumstances and use cases, taking into account the issues highlighted in the previous section that apply to you. The cost-benefit analysis will doubtless change over time. We can’t predict the future of the AI industry or hardware prices even in the short term.

Introducing Jina On-Prem

For users who can benefit from local AI services, we’re introducing Jina On-Prem, a fully self-contained installation suite for Jina AI’s high-performance models.

Jina AI’s models match the accuracy of embedding models many times their size, reducing compute costs, memory footprints, and hardware requirements. This makes them an ideal choice for users who want or need to keep their AI on-premises. Commercial licenses are available with scalable, proportionately priced solutions for use cases of all sizes.

What API schemas does Jina On-Prem support?

  • Available as a complete collection of dependencies for local installation or as a Docker container that you can install and run in minutes.
  • Jina On-Prem installations do not call out to outside systems.
    • No call to Hugging Face Hub or any model registry (HF_HUB_OFFLINE=1 and TRANSFORMERS_OFFLINE=1 are baked in).
    • There’s no license server.
    • There are no telemetry or logging endpoints.
  • Supports both CPU and GPU hardware, with GPU autodetection.
  • All 28 Jina AI models available, including the latest jina-embeddings-v5-omni multimodal embedding models and jina-reranker-v3.
  • Access via standard AI API schemas: Jina API, OpenAI, Cohere, Voyage AI, and Gemini. Jina On-Prem is a drop-in solution for applications built on those schemas.
  • Drop-in replacement for models served by the EIS. Jina On-Prem integrates directly with air-gapped Elastic deployments.

Hardware requirements for Jina AI on-prem models

The hardware requirements vary for different Jina models. The table below shows the recommendations for the most recent models using GPU settings. You don’t need anything more powerful than an NVIDIA L4 GPU, although an A100 is recommended for the v5 embedding models. Our latest embedding model currently requires a minimum of 8 GB of VRAM.

ModelMinimum VRAMRecommended GPU
jina-embeddings-v5-text-nano2 GBT4 / L4
jina-embeddings-v5-text-small3 GBL4 / A10G
jina-embeddings-v5-omni-small8 GBL4 / A10G / A100
jina-reranker-v33 GBL4
jina-clip-v24 GBL4
jina-code-embeddings-1.5b4 GBL4
ReaderLM-v24 GBL4

If you use more than one model at a time, the VRAM requirements will increase. Please see the Sizing and Hardware page for more information.

How to install Jina On-Prem with Docker

The quickest way to get started is to install Docker (if you haven’t already) and follow the instructions on the Jina On-Prem Quick Start page.

There are pre-composed Docker containers for all 28 Jina models. Download one and transfer it to your installation target, and you can have Jina AI models running in under five minutes.

For multimodal or custom builds, or to download the complete dependency set for installation outside of a container, follow the steps outlined in the bundling guide.

You’ll need a GitHub account and access token to download Jina On-Prem. To create a free account, go to https://github.com/signup. To generate or manage your access tokens, follow the instructions in the GitHub documentation.

Your Jina On-Prem installation supports all Jina API and EIS functionality and embedding generation via OpenAI, Cohere, Voyage AI, and Gemini APIs, so it can integrate into preexisting applications using standard interfaces. See the API documentation for more information.

Jina models, including models installed with Jina On-Prem, are available on various licensing terms, with the latest models free for noncommercial use under a CC BY-NC 4.0 license. To license Jina On-Prem for commercial use, please contact Elastic Sales.

¿Te ha sido útil este contenido?

No es útil

Algo útil

Muy útil

Contenido relacionado

¿Estás listo para crear experiencias de búsqueda de última generación?

No se logra una búsqueda suficientemente avanzada con los esfuerzos de uno. Elasticsearch está impulsado por científicos de datos, operaciones de ML, ingenieros y muchos más que son tan apasionados por la búsqueda como tú. Conectemos y trabajemos juntos para crear la experiencia mágica de búsqueda que te dará los resultados que deseas.

Pruébalo tú mismo