{"id":689,"date":"2024-03-14T21:56:15","date_gmt":"2024-03-14T10:56:15","guid":{"rendered":"https:\/\/www.samontab.com\/web\/?p=689"},"modified":"2024-03-14T21:56:16","modified_gmt":"2024-03-14T10:56:16","slug":"openvino-performance-for-state-of-the-art-real-time-monocular-depth-estimation","status":"publish","type":"post","link":"https:\/\/www.samontab.com\/web\/2024\/03\/openvino-performance-for-state-of-the-art-real-time-monocular-depth-estimation\/","title":{"rendered":"OpenVINO performance for state of the art real-time monocular depth estimation"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Recently, an interesting paper was accepted to CVPR 2024, <a href=\"https:\/\/depth-anything.github.io\/\">Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data<\/a>. The authors made publicly available their pre-trained models in three sizes: Depth-Anything-Small, Depth-Anything-Base, and Depth-Anything-Large. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I wanted to have an idea of how fast these models can run on a CPU these days. Since I am interested in real-time operation, I started with the smallest model to get an idea of how well it performed and how fast it can run the inference in my laptop:<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img data-dominant-color=\"4d2f2a\" data-has-transparency=\"false\" style=\"--dominant-color: #4d2f2a;\" loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"346\" src=\"https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/bir.png\" alt=\"\" class=\"wp-image-694 not-transparent\" srcset=\"https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/bir.png 1024w, https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/bir-300x101.png 300w, https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/bir-768x260.png 768w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<figure class=\"wp-block-gallery has-nested-images columns-default is-cropped wp-block-gallery-1 is-layout-flex wp-block-gallery-is-layout-flex\">\n<figure class=\"wp-block-image size-large\"><img data-dominant-color=\"754e50\" data-has-transparency=\"false\" style=\"--dominant-color: #754e50;\" loading=\"lazy\" decoding=\"async\" width=\"1020\" height=\"768\" data-id=\"695\" src=\"https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/stree.png\" alt=\"\" class=\"wp-image-695 not-transparent\" srcset=\"https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/stree.png 1020w, https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/stree-300x226.png 300w, https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/stree-768x578.png 768w\" sizes=\"auto, (max-width: 1020px) 100vw, 1020px\" \/><\/figure>\n<\/figure>\n\n\n\n<figure class=\"wp-block-image size-full\"><img data-dominant-color=\"785c61\" data-has-transparency=\"false\" style=\"--dominant-color: #785c61;\" loading=\"lazy\" decoding=\"async\" width=\"1017\" height=\"768\" src=\"https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/gira.png\" alt=\"\" class=\"wp-image-697 not-transparent\" srcset=\"https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/gira.png 1017w, https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/gira-300x227.png 300w, https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/gira-768x580.png 768w\" sizes=\"auto, (max-width: 1017px) 100vw, 1017px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">You can see that the smallest model still performs relatively well. In terms of inference time, it took on average 1.02 seconds per image on my CPU, and the size of this PyTorch model, depth_anything_vits14.pth, is 95MB. It&#8217;s fast and small, but not really ideal for real-time applications. Let&#8217;s see if we can do better.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">OpenVINO uses their own Intermediate Representation format, IR, which is designed to be optimised for inference. Furthermore, you can then generate a quantised model from the already OpenVINO optimised model, getting even better inference times and a smaller model.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Here is a table with the different results on my machine:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table><tbody><tr><td><strong>Model Name<\/strong><\/td><td><strong>Inference speed (FPS)<\/strong><\/td><td><strong>Model Size (MB)<\/strong><\/td><\/tr><tr><td>depth_anything_vits14.pth (Original PyTorch model)<\/td><td>~1<\/td><td>95<\/td><\/tr><tr><td>depth_anything_vits14.(bin+xml) (OpenVINO IR)<\/td><td>7.65<\/td><td>47.11<\/td><\/tr><tr><td>depth_anything_vits14_int8.(bin+xml) (IR Quantised)<\/td><td><strong>9.74<\/strong><\/td><td><strong>24.27<\/strong><\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">You can clearly see that there is a massive increase in performance when you use the quantised OpenVINO IR model compared to the original PyTorch model. This allows real time operation on a laptop. And for reference, here is the output of the quantised model:<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img data-dominant-color=\"43172a\" data-has-transparency=\"false\" style=\"--dominant-color: #43172a;\" loading=\"lazy\" decoding=\"async\" width=\"683\" height=\"1024\" src=\"https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/image-683x1024.png\" alt=\"\" class=\"wp-image-699 not-transparent\" srcset=\"https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/image-683x1024.png 683w, https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/image-200x300.png 200w, https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/image-768x1152.png 768w, https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/image-1024x1536.png 1024w, https:\/\/www.samontab.com\/web\/wp-content\/uploads\/2024\/03\/image-1365x2048.png 1365w\" sizes=\"auto, (max-width: 683px) 100vw, 683px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The output is still reasonably good, with the nice bonus that it can be used in real time, plus the file size is about one quarter of the original. All thanks to OpenVINO&#8217;s great ability to optimise the inference pipeline!<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Recently, an interesting paper was accepted to CVPR 2024, Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. The authors made publicly available their pre-trained models in three sizes: Depth-Anything-Small, Depth-Anything-Base, and Depth-Anything-Large. I wanted to have an idea of how fast these models can run on a CPU these days. Since I am interested [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[29,21,24],"tags":[46,51,25],"class_list":["post-689","post","type-post","status-publish","format-standard","hentry","category-computer-vision","category-open-source","category-openvino","tag-depth","tag-monocular","tag-openvino"],"_links":{"self":[{"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/posts\/689","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/comments?post=689"}],"version-history":[{"count":0,"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/posts\/689\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/media?parent=689"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/categories?post=689"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.samontab.com\/web\/wp-json\/wp\/v2\/tags?post=689"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}