dvar

DVAR MODELS / OPEN-WEIGHT LIBRARY

Choose the smallest model that can do the job.

A practical catalog for Indian enterprise teams. Start with capability and economics, then validate the model on your own workload before you commit.

MODEL ROUTERQUALITY / COST
WORKFLOW FIT82
UNIT ECONOMICS94
CONTROL100

A model is only “best” after it wins on your data, latency, and budget.

DVAR MODEL LIBRARY / WORKING CATALOG

Open models, with the details that matter in a buying conversation.

Our internal catalog organizes open-weight model families by capability, context, license, and deployment fit. We keep it current as model cards and production checks change.

21 models tracked6 capability areas3 Dvar lanesUpdated August 2026
Get a workload-specific shortlist
MODELCAPABILITYMODEL DETAILSDVAR FIT
Llama 3.3 70B InstructMeta Llama · model card
Chat & reasoning70B128K context · Llama communityBalancedReliable general chat, extraction, multilingual work
Llama 3.1 8B InstructMeta Llama · model card
Chat & reasoning8B128K context · Llama communityEcoLow-latency chat, classification, high-volume routing
Llama 4 ScoutMeta Llama · model card
VisionMoELong context context · Llama communityBalancedMultimodal understanding and document-heavy workflows
Qwen3 8BQwen · model card
Chat & reasoning8BLong context context · Apache 2.0EcoFast reasoning, structured extraction, multilingual tasks
Qwen3 32BQwen · model card
Chat & reasoning32BLong context context · Apache 2.0BalancedReasoning, tool use, enterprise assistants
Qwen3 235B A22BQwen · model card
Chat & reasoning235B MoE / 22B activeLong context context · Apache 2.0ProComplex reasoning and high-value agentic work
DeepSeek-V3.1DeepSeek · model card
Chat & reasoningMoELong context context · MITProGeneral reasoning, coding, and cost-efficient frontier work
DeepSeek-R1DeepSeek · model card
Chat & reasoning671B MoE128K context · MITProDeliberate reasoning, math, research, difficult decisions
Kimi K2Moonshot AI · model card
Code & agents1T MoELong context context · Modified MITProLong-horizon coding, tool use, research agents
GLM-4.5 AirZhipu AI · model card
Code & agentsMoELong context context · GLM licenseBalancedFast agent loops, coding, function calling
Mistral Small 3.1Mistral AI · model card
Vision24B128K context · Apache 2.0BalancedCompact multimodal assistant, extraction, on-premise fit
Gemma 3 4BGoogle Gemma · model card
Vision4B128K context · Gemma termsEcoSmall multimodal workflows and edge-friendly assistants
Gemma 3 27BGoogle Gemma · model card
Chat & reasoning27B128K context · Gemma termsBalancedStrong quality in a smaller deployment footprint
GPT-OSS 20BOpenAI open-weight · model card
Chat & reasoning21B MoE128K context · Apache 2.0EcoOpen-weight reasoning and low-cost production tasks
GPT-OSS 120BOpenAI open-weight · model card
Code & agents117B MoE128K context · Apache 2.0ProTool use, coding, reasoning, high-quality open deployment
Qwen3 Embedding 8BQwen · model card
Embeddings8B32K context · Apache 2.0BalancedMultilingual retrieval and semantic search
BGE-M3BAAI · model card
Embeddings568M8K context · MITEcoHybrid multilingual retrieval and reranking
Nomic Embed v1.5Nomic AI · model card
Embeddings137M8K context · Apache 2.0EcoAffordable semantic search, clustering, and RAG
Whisper Large v3OpenAI open-weight · model card
Speech1.5BAudio context · MITBalancedMultilingual transcription and meeting intelligence
FLUX.1 [dev]Black Forest Labs · model card
Image12BImage context · FLUX termsBalancedHigh-quality image generation and creative prototyping
Stable Diffusion XLStability AI · model card
Image3.5BImage context · CreativeMLEcoImage generation with a broad ecosystem and controls

Model names, context windows, licenses, and deployment availability change. This page is Dvar’s working catalog and decision aid, not a promise that every model is available in every deployment mode. Dvar benchmarks the selected model on your workload before recommending production use.

HOW TO CHOOSE

Quality is not one number.
Use the workload.

Choose Eco when volume matters first

Use smaller models for classification, extraction, routing, and high-volume employee work where a 10% quality gain may not justify a 4× token bill.

Choose Balanced when the task is mixed

Use a capable 8B–32B model for everyday assistants, structured outputs, and production workloads with a real latency budget.

Choose Pro when the decision is expensive

Use larger reasoning or agentic models for complex research, code, and workflow design, then decide whether a tuned smaller model can take over the repeatable core.

MODEL EVALUATION

Bring us your benchmark.
We will bring the shortlist.

Share 20–50 representative examples and your latency and budget constraints. We will return a decision-ready comparison, not a leaderboard tour.

Contact Us