LLM Providers

vLLM Setup

Connect MLJAR Studio to a vLLM server on your workstation or on a separate machine. The server runs the model; Studio connects to its API.

Before you start

Activate a lifetime MLJAR Studio license to configure your own AI provider. Open AI Provider Settings in MLJAR Studio to choose a provider.

1. Start the model server

Follow the vLLM installation instructions for your hardware. Choose a supported chat model and start the server. Replace YOUR_MODEL_ID with its model identifier:

vllm serve YOUR_MODEL_ID --host 127.0.0.1 --port 8000

Wait for model loading to finish. See the vLLM server documentation for chat templates and server options. If the server uses --api-key, use that key in Studio.

2. Connect MLJAR Studio

  1. In AI Provider Settings, select vLLM from Provider.
  2. Set Base URL to http://127.0.0.1:8000/v1 for the default local setup. Use the actual host and port if you changed them. Keep the /v1 suffix.
  3. Leave vLLM API key (optional) empty for a server without authentication, or enter the key required by your server.
  4. Wait for models to load automatically, or click Refresh models. Choose your chat model from Model.
  5. Click Test connection, then Save provider once the test succeeds. Changing settings requires another test.
  6. Confirm the success message and the active provider in the sidebar.

Test a notebook request

After saving, send a short request in your notebook. Test connection checks access to the model list and the selected model; a successful test does not guarantee that every model supports the chat or tool features your workflow needs.

Where requests run

The default loopback address connects to the server on the machine running the MLJAR notebook backend. A remote server URL sends prompts and notebook context to that server. Use HTTPS for remote connections. For local inference, keep both the endpoint and model execution on your machine.

Troubleshooting

  • If connection is refused, check that vLLM finished starting and is listening on the configured port.
  • If generation fails after a successful connection test, check that the model supports chat and has the required chat template.
  • If the server runs out of memory, check its logs and choose a model and context size that fit the server hardware.

Related guides

All LLM providers · Local vs Cloud LLMs · LLM Setup Troubleshooting

« Previous
Jan Setup