modelparams.dev
NVIDIA Chat Completions 5 params

NVIDIA Llama 3.2 90B Vision Instruct API parameters

These are the API parameters modelparams.dev tracks for NVIDIA Llama 3.2 90B Vision Instruct on Chat Completions with an API key — the settings you send in a request. Each row gives the type, default, valid range or values, and the conditions that gate it. It's the same data the JSON API serves.

After the parameter count instead — how many weights Llama 3.2 90B Vision Instruct has? That's a different number, and we don't track it. Here's the difference.

Length 1 param
Parameter Type Default Description Condition
Max tokens
max_tokens
integer (1…16384) 4096 Maximum number of tokens to generate. Generation stops when this limit is reached. —
Sampling 4 params
Parameter Type Default Description Condition
Temperature
temperature
number (0…1) 0.6 Controls randomness. Lower values make outputs more focused; higher values make them more varied. Not recommended to modify both temperature and top_p in the same call. —
Top P
top_p
number (-∞…1) 0.95 Controls nucleus sampling by limiting generation to tokens within the selected cumulative probability. Not recommended to modify both temperature and top_p in the same call. —
Frequency penalty
frequency_penalty
number (-2…2) 0 Penalizes new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim. —
Presence penalty
presence_penalty
number (-2…2) 0 Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics. —

NVIDIA Llama 3.2 90B Vision Instruct API parameters in brief

NVIDIA Llama 3.2 90B Vision Instruct on Chat Completions documents 5 API parameters, grouped by what they control:

Frequently asked questions

Which API parameters does NVIDIA Llama 3.2 90B Vision Instruct support?
NVIDIA Llama 3.2 90B Vision Instruct accepts 5 API parameters in the request body: temperature, top_p, max_tokens, frequency_penalty, presence_penalty.
What is the default temperature for NVIDIA Llama 3.2 90B Vision Instruct?
The default temperature for NVIDIA Llama 3.2 90B Vision Instruct is 0.6, within a valid range of 0 to 1.
What is the default top_p for NVIDIA Llama 3.2 90B Vision Instruct?
The default top_p for NVIDIA Llama 3.2 90B Vision Instruct is 0.95, with a maximum of 1.
What is the default max_tokens for NVIDIA Llama 3.2 90B Vision Instruct?
The default max_tokens for NVIDIA Llama 3.2 90B Vision Instruct is 4096, within a valid range of 1 to 16384.

Resources

All NVIDIA models Glossary Full catalog

Llama 3.2 90B Vision Instruct — JSON

The full model definition as served by the API. Copy it or open the endpoint directly.

{
  "$schema": "https://modelparams.dev/api/v1/schema.json",
  "provider": "nvidia",
  "authType": "api_key",
  "apiSurface": "openai-chat-completions",
  "model": "llama-3.2-90b-vision-instruct",
  "wireId": "meta/llama-3.2-90b-vision-instruct",
  "params": [
    {
      "path": "temperature",
      "label": "Temperature",
      "description": "Controls randomness. Lower values make outputs more focused; higher values make them more varied. Not recommended to modify both temperature and top_p in the same call.",
      "group": "sampling",
      "type": "number",
      "default": 0.6,
      "range": {
        "min": 0,
        "max": 1
      }
    },
    {
      "path": "top_p",
      "label": "Top P",
      "description": "Controls nucleus sampling by limiting generation to tokens within the selected cumulative probability. Not recommended to modify both temperature and top_p in the same call.",
      "group": "sampling",
      "type": "number",
      "default": 0.95,
      "range": {
        "max": 1
      }
    },
    {
      "path": "max_tokens",
      "label": "Max tokens",
      "description": "Maximum number of tokens to generate. Generation stops when this limit is reached.",
      "group": "generation_length",
      "type": "integer",
      "default": 4096,
      "range": {
        "min": 1,
        "max": 16384
      }
    },
    {
      "path": "frequency_penalty",
      "label": "Frequency penalty",
      "description": "Penalizes new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim.",
      "group": "sampling",
      "type": "number",
      "default": 0,
      "range": {
        "min": -2,
        "max": 2
      }
    },
    {
      "path": "presence_penalty",
      "label": "Presence penalty",
      "description": "Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.",
      "group": "sampling",
      "type": "number",
      "default": 0,
      "range": {
        "min": -2,
        "max": 2
      }
    }
  ]
}

Other NVIDIA models

DeepSeek v4 Flash 0731 7 params View DeepSeek v4 Pro 0813 7 params View Diffusiongemma 26B A4B IT 5 params View Gemma 4 31B IT 7 params View GLiNER PII 4 params View GLM-5.3 Flash 7 params View GPT-OSS 120B 7 params View GPT-OSS 20B 7 params View Kimi K3 7 params View Laguna Xs 2.1 7 params View Llama 3.1 NemoGuard 8B Topic Control 6 params View Llama 3.1 Nemotron Nano 8B v1 7 params View Llama 3.1 Nemotron Safety Guard 8B v3 1 param View Llama 3.1 Nemotron Ultra 253B v1 7 params View Llama 3.2 11B Vision Instruct 5 params View Llama 3.3 Nemotron Super 49B v1 7 params View Llama 3.3 Nemotron Super 49B v1.5 7 params View MiniMax M3 7 params View Mistral Nemotron 5 params View Muse Glimmer 30B 7 params View NemoGuard Jailbreak Detect 0 params View Nemotron 3 Nano 30B A3B 5 params View Nemotron 3 Nano Omni 30B A3B Reasoning 3 params View Nemotron 3 Super 120B A12B 7 params View Nemotron 3 Ultra Subscription 6 params View Nemotron 3 Ultra 550B A55B 7 params View Nemotron 3.5 Lightning 30B A3B 5 params View Nemotron Content Safety Reasoning 4B 5 params View Nemotron Mini 4B Instruct 6 params View Riva Translate 4B Instruct v1.1 6 params View USDCode Llama 3.1 70B Instruct 4 params View

How to use

Building with an AI agent? Hit Copy to grab this whole guide as Markdown and paste it in — or point your agent straight at /llms.txt.

modelparams.dev is an open, community-maintained catalog of model parameters. Each entry shows the knobs you can turn — type, default, range, and the conditions that gate it.

The same model accessed via an API key and via a subscription usually exposes a different set of parameters. We list both as separate entries so the data stays honest.

Catalog API

The full catalog is static JSON, CORS-enabled, served from the edge.

curl https://modelparams.dev/api/v1/models.json

Each entry is keyed by provider/model for API-key variants; subscription variants append -subscription.

If you only need the params for one model contract, use the providerless endpoint. Subscription contracts are model slugs with -subscription.

curl https://modelparams.dev/api/v1/models/openai/gpt-5.5.json
curl https://modelparams.dev/api/v1/models/openai/gpt-5.5-subscription.json

Single model

curl https://modelparams.dev/api/v1/models/anthropic/claude-opus-4-7.json
curl https://modelparams.dev/api/v1/models/anthropic/claude-opus-4-7-subscription.json

JSON Schema

Every entry validates against a JSON Schema you can use in your editor or pipeline.

curl https://modelparams.dev/api/v1/schema.json

Add this header to any YAML you author for autocomplete in VS Code:

# yaml-language-server: $schema=https://modelparams.dev/api/v1/schema.json

Logos

Provider logos are available at /assets/logos/{provider}.svg where {provider} is the provider slug. They use currentColor so they inherit your text color.

curl https://modelparams.dev/assets/logos/anthropic.svg

Logos are sourced from the models.dev repo (MIT) and used under nominative fair use.

Contribute

The data lives in YAML under models/{provider}/{model}-{auth}.yaml in the GitHub repo. Open a PR; CI validates against the schema and rebuilds.

Edit on GitHub MIT licensed