Large memory GPU instances for deploying and running large language models like Llama, ChatGLM, and more.