LLM Providers

llama.cpp Setup

Use llama-server from llama.cpp to serve a local GGUF model for MLJAR Studio. Choose a chat-capable model that fits your available memory.

Before you start

Activate a lifetime MLJAR Studio license to configure your own AI provider. Open AI Provider Settings in MLJAR Studio to choose a provider.

1. Start the model server

Install or build llama.cpp using its official repository. Download a compatible GGUF chat model, then start llama-server. Replace the example path with your model file:

llama-server -m /path/to/model.gguf --host 127.0.0.1 --port 8080

Keep the process running. The llama.cpp server guide covers installation options, model loading, authentication, and chat templates.

2. Connect MLJAR Studio

  1. In AI Provider Settings, select llama.cpp from Provider.
  2. Set Base URL to http://127.0.0.1:8080/v1 for the default local setup. Use the actual host and port if you changed them. Keep the /v1 suffix.
  3. Leave llama.cpp API key (optional) empty unless your server requires authentication; otherwise enter its key.
  4. Wait for models to load automatically, or click Refresh models. Choose your chat model from Model.
  5. Click Test connection, then Save provider once the test succeeds. Changing settings requires another test.
  6. Confirm the success message and the active provider in the sidebar.

Test a notebook request

After saving, send a short request in your notebook. Test connection checks access to the model list and the selected model; a successful test does not guarantee that every model supports the chat or tool features your workflow needs.

Where requests run

The default loopback address connects to the server on the machine running the MLJAR notebook backend. A remote server URL sends prompts and notebook context to that server. Use HTTPS for remote connections. For local inference, keep both the endpoint and model execution on your machine.

Troubleshooting

  • If the command is not found, locate the llama-server executable from your installation or build.
  • If the model does not load, check the GGUF file path and available memory in the server logs.
  • Select the model identifier returned by the server, which may differ from the filename. Check the chat template if requests fail after connecting.

Related guides

All LLM providers · Local vs Cloud LLMs · LLM Setup Troubleshooting

« Previous
vLLM Setup