{"id":707,"date":"2024-09-25T23:39:09","date_gmt":"2024-09-25T13:39:09","guid":{"rendered":"https:\/\/www.samontab.com\/web\/?p=707"},"modified":"2024-09-26T19:01:29","modified_gmt":"2024-09-26T09:01:29","slug":"how-to-run-llama-3-1-locally-on-mac-and-serve-it-to-a-local-linux-laptop-to-use-with-zed","status":"publish","type":"post","link":"https:\/\/www.samontab.com\/web\/2024\/09\/how-to-run-llama-3-1-locally-on-mac-and-serve-it-to-a-local-linux-laptop-to-use-with-zed\/","title":{"rendered":"How to run Llama 3.2 locally on Mac and serve it to a local Linux laptop to use with Zed"},"content":{"rendered":"\n<figure class=\"wp-block-image size-full is-resized\"><img data-dominant-color=\"292d34\" data-has-transparency=\"false\" loading=\"lazy\" decoding=\"async\" width=\"761\" height=\"878\" src=\"https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/09\/llama_3.2.gif\" alt=\"\" class=\"wp-image-713 not-transparent\" style=\"--dominant-color: #292d34; width:410px;height:auto\"\/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>UPDATE<\/strong>: I wrote this post for Llama3.1, but just after I published it, Llama3.2 was released. It comes with similar performance but faster inference as it\u2019s a distilled model(~2.6 times faster than Llama3.1 in a quick test I did). I updated this post to use 3.2 but it should work with any other version as well.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/zed.dev\/\" target=\"_blank\" rel=\"noreferrer noopener\">Zed<\/a> is a great editor that supports AI assistants. In this post I will explain how you can share one Llama model you have running in a Mac between other computers in your local network for privacy and cost efficiency. Also, fans might get loud if you run Llama directly on the laptop you are using Zed as well.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Since I&#8217;ve found that Apple silicon (M1, M2, etc) is quite good at running these models, I will assume the model will be run in that computer. The default LLama3.2:3b works fine on a Mac Mini M1 with 16GB. If you have 8GB you might need to use simpler models.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The first step is to install <a href=\"https:\/\/ollama.com\/\" target=\"_blank\" rel=\"noreferrer noopener\">ollama<\/a> in your Mac. Just follow the instruction on the website. You will end up with a little llama icon on menu bar(top right). We now need to install a model, Llama3.2 in particular. In the terminal run this:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code lang=\"bash\" class=\"language-bash\">ollama run llama3.2<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">This will try to run llama3.2 and since it is not yet installed, it will fetch the latest model for you. After it downloads and runs the model, simply type something to test it all works. It should look like this:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code lang=\"bash\" class=\"language-bash\">&gt;&gt;&gt; Tell me a joke\nHere's one:\n\nWhat do you call a fake noodle?\n\nAn impasta.\n\n&gt;&gt;&gt; Send a message (\/? for help)<\/code><\/pre>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\"><\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">Now that it works locally we want to make it available for other computers in the local network. Open the terminal and run this command:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code lang=\"bash\" class=\"language-bash\">launchctl setenv OLLAMA_HOST 0.0.0.0:11434<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Now click on the icon and exit ollama. Then start ollama again. It should now be ready to accept connections from other computers in your network. To check connectivity, go to a Linux computer in your network and open a terminal. Run the following:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code lang=\"bash\" class=\"language-bash\">curl http:\/\/your_mac.local:11434\/api\/generate -d '{\n&nbsp;\"model\": \"llama3.2\",\n&nbsp;\"prompt\": \"Tell me a joke\",\n&nbsp;\"options\": {\n&nbsp;&nbsp;&nbsp;\"num_ctx\": 4096\n&nbsp;}\n}'<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Make sure you see a correct response before continuing. Now let&#8217;s configure Zed to use it. Open settings (CTRL-SHIFT-P and write open settings). Add these settings there (plus any others you already have, it&#8217;s a json file):<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code lang=\"javascript\" class=\"language-javascript\">{\n  \"language_models\": {\n    \"ollama\": {\n      \"api_url\": \"http:\/\/your_mac.local:11434\",\n      \"low_speed_timeout_in_seconds\": 120,\n      \"keep_alive\": \"120s\",\n      \"available_models\": [\n        {\n          \"provider\": \"ollama\",\n          \"name\": \"llama3.2:latest\",\n          \"max_tokens\": 16384\n        }\n      ]\n    }\n  },\n  \"assistant\": {\n    \"default_model\": {\n      \"provider\": \"ollama\",\n      \"model\": \"llama3.2:latest\"\n    },\n    \"version\": \"2\"\n  }\n}<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">It should now be configured. You can go to the Assistant Panel (CTRL+?) and ask whatever you want there. You can add context as well with <strong>\/tab<\/strong> and others.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That&#8217;s it, now you have a shared Llama 3.2 model running in one computer, while using it on another computer in your network, privately and free, integrated into a great text editor.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Run Llama3.1 on a mac and share it to a Linux laptop on your local network using Zed.<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[97,21,4,98],"tags":[104,102,100,105,101,103,106,99],"class_list":["post-707","post","type-post","status-publish","format-standard","hentry","category-ai","category-open-source","category-programming","category-zed","tag-free","tag-linux","tag-llama","tag-local","tag-mac","tag-private","tag-shared","tag-zed"],"_links":{"self":[{"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/posts\/707","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/comments?post=707"}],"version-history":[{"count":0,"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/posts\/707\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/media?parent=707"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/categories?post=707"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/tags?post=707"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}