The Algorithmic #1: Model Rundown

Welcome to The Algorithmic

This is the inaugural blog post in our new series, The Algorithmic: Your Look at AI’s Edge. In this blog series, we’ll be bringing you new developments in the world of AI, as well as tutorials and insights. This week was a landmark week for AI, with a staggering number of new state-of-the-art models being released. We’ll be kicking things off with a rundown of the AI models developed in the past week, still hot off the GPUs. Without further ado, let’s dig in!

GPT-4

This week, OpenAI held their DevDay. One of their major announcements was GPT-4 Turbo, the latest in the GPT series that powers their ChatGPT service. GPT-4 Turbo boasts a 128k context window, meaning users can submit prompts of over 300 pages of text! Need to get up-to-speed quickly for a book club, or ask a question about an entire codebase? GPT-4 Turbo can do that.

It has also been trained on more data to give it knowledge of world events up through April 2023 and has been optimized to compose results much faster. Another benefit of the optimizations made over GPT-4 is that API prices have been slashed (about half the price for input tokens and a third of the price for output tokens).

If all that wasn’t enough, GPT-4 Turbo has enhanced features suited to developers, such as function calling, which allows it to interface with functions in a user-defined app or external APIs. It’s also had an upgrade to its abilities to respond in structured formats, like JSON and XML.

As an added bonus, GPT-4 Turbo allows users to specify a random number seed for reproducible results (which is invaluable for debugging code that uses GPT-4 Turbo over the API), as well as the option to output log probabilities for the most likely tokens in its response (allowing developers to take GPT outputs and optimize/customize them to their heart’s content).

Grok-1

OpenAI isn’t the only one with big announcements this week. Days before, X (formerly Twitter) announced Grok-1, the first AI model from Elon Musk’s venture xAI. Grok-1 is a state-of-the-art language model that has shown impressive results on various machine learning benchmarks, surpassing all other models in its compute class, including GPT-3.5 and Inflection-1. However, it still lags behind the industry frontrunner, GPT-4.

The model, which powers the Grok chatbot client, was developed with significant improvements in reasoning and coding capabilities.

The Grok-1 language model has 33 billion parameters with a context window of 8192 tokens. Impressively, it was developed in a relatively short amount of time by industry standards – just two months.

In true Elon Musk style, it’s designed to answer questions and suggest what questions to ask, all with a dash of humor inspired by “The Hitchhiker’s Guide to the Galaxy.”

OpenChat-3.5

OpenChat-3.5 is an innovative open-source LLM that’s been fine-tuned with C-RLFT, a strategy inspired by offline reinforcement learning, and learns from mixed-quality data without preference labels. Despite its simple approach, OpenChat-3.5 delivers exceptional performance on par with ChatGPT, even with a 7B model!

Nous-Yarn-Mistral-7b-128k

Nous-Yarn-Mistral-7b-128k is a model from the popular Mistral-7b series, and is known for its impressive 128k context window. This model stands out in the AI landscape for its ability to handle extensive context, making it a powerful tool for tasks that require a deep understanding of large amounts of text.

OpenHermes 2.5 – Mistral-7b

OpenHermes 2.5 is a fine-tuned version of the Mistral-7b model, specifically optimized for enhanced coding ability. Surprisingly, it also performs significantly better on several non-code benchmarks, demonstrating its versatility and robustness. As part of the Mistral-7b family, OpenHermes 2.5 is a testament to the power of fine-tuning in enhancing the capabilities of large language models.

Embed v3 (Cohere)

Cohere’s Embed v3 is the latest and most advanced text embedding model, offering state-of-the-art performance on trusted benchmarks such as the Massive Text Embedding Benchmark (MTEB) and the Benchmark for Evaluating Information Retrieval (BEIR). One of the key improvements in Embed v3 is its ability to evaluate how well a query matches a document’s topic and assess the overall quality of the content. This makes it particularly effective when dealing with noisy datasets.

HelixNet

HelixNet is a large language model architecture consisting of three Mistral-7b LLMs. The three sub-nets are called the actor, critic, and regenerator. Given a user prompt, the actor produces an initial response. Then, the critic takes the prompt and response as inputs and provides a critique. Finally, the regenerator takes the prompt, initial response, and critique, then produces a refined response.

One of the great things about this approach is you don’t have to use any particular LLM. This opens up possibilities for task-specific HelixNet ensembles.

DeepSeek Coder

DeepSeek Coder, the AI that’s more into coding than a caffeine-fueled developer during a hackathon, is a series of code language models trained on a whopping 2 trillion tokens, with a composition of 87% code and 13% natural language. It’s available in sizes from 1.3B to 33B parameters. Coder is making waves on benchmarks, boasting state-of-the-art performance on benchmarks like HumanEval, MultiPL-E, MBPP, DS-1000, and APPS. It’s got a window size of 16K and a fill-in-the-blank task, supporting project-level code completion and infilling tasks. It’s also multilingual, so whether you’re coding in English, Chinese, or one of over 80 other languages, DeepSeek Coder has got your back.

Whisper-Large-3 and Distil-Whisper

Whisper-Large-3, the latest iteration of OpenAI’s Whisper, is a speech recognition model trained on a staggering 1 million hours of weakly labeled audio and 4 million hours of pseudo-labeled audio. This model, a veritable titan of transcription, is a testament to the power of large-scale training data and the advancements in automatic speech recognition (ASR) technology.

On the other hand, Distil-Whisper is a smaller, faster, and more efficient variant of the Whisper model, distilled using a large-scale pseudo-labeling technique. This model is a marvel of model compression, maintaining the robustness of the original Whisper model while being 5.8 times faster and having 51% fewer parameters. It performs within a 1% word error rate (WER) on out-of-distribution test data in a zero-shot transfer setting, making it a compact powerhouse in the realm of ASR.

DALL-E 3 VAE

DALL-E 3 VAE is the visual post-processing component in OpenAI’s celebrated DALL-E 3 text-to-image model. Earlier this week, OpenAI released the DALL-E 3 VAE to the public, making it available for hacker-artists to code into their digital art pipelines.

As a VAE, it can plug directly into the Stable Diffusion pipeline for an instant boost in visual quality. Be on the lookout for the whole Stable Diffusion world to take this little gem and bring the world some crazy new creations.

jina-embeddings-v2

Jina AI has released the second version of their text embeddings model. Not only does it boast impressive performance on benchmarks, it has a context length of 8192. This is fantastic news for anyone building a Retrieval-Augmented Generation (RAG) system, because it allows for direct storage of larger documents (careful though, it’s been observed that bigger chunks don’t necessarily mean better results). Currently, only the Base (137 million parameters) and Small (33 million parameters) versions of the model are available, though the Large variant (435 million parameters) should be released soon.

Summary

Whether you use AI assistants, produce AI art or build production systems using AI, this has been an exciting week. While OpenAI and xAI dominated the news cycle with their announcements, this has been a great week for developers working on RAG systems, with context windows extending to extraordinary lengths. There’s a fantastic selection of new LLMs to try out, a new VAE for your Stable Diffusion pipelines, as well as state-of-the-art and speed in speech-to-text.

Thanks for stopping in to catch up with us on AI. We’ll be back soon with more models, news and tutorials.

Secret Link