The Algorithmic #3: Hot off the Compiler – The Latest AI Apps and Libraries

In this week’s post we’ll be looking at everything new in ML/AI space. Without further ado…

Generative AI

Stable Diffusion

Video Generation

Generative AI:Video Generation

Right now, text-to-video and image-to-video are getting a lot of attention. Whether fueled by blazingly fast diffusion models (Latent Consistency Models, SDXL Turbo) with plug-ins or native video models like Stable Video Diffusion, Runway or Pika, generative video is moving fast.

  • ComfyUI made a major update, allowing users to run Stable Video Diffusion on just 8GB of vRAM at 25+ frames per second. For all of you who love wiring together ComfyUI workflows, this is a big one.
  • Over the past few days, MagicAnimate has been going viral, turning a sizable fraction of the internet into TikTok-style dance videos.
  • A new app, controlGIF, just dropped. It allows users to combine ControlNets with AnimateDiff to generate short videos using Stable Diffusion v1.5.
  • Want to try Runway Labs’ motion brush, but want to stay open source? Now there’s a new app that does just that!
  • Pika Labs announced Pika 1.0.
  • Decohere released a preview of their AI video generator.
  • Runway Studios is partnering with Getty Images, giving them a massive new source of data to create a new suite of generative video models.
  • As an added bonus, there’s a new ComfyUI implementation of Robust Video Matting.

ComfyUI Workflow Drops

It’s not Christmas yet, but it sure feels like it. Within the span of a few days, no fewer than three top-tier AI artists have released collections of their ComfyUI workflows. There’s enough material here to keep an aspiring AI artist tinkering for months!

ZipLoRA in Less Than A Week

Less than a week after the ZipLora paper came out, there was already a PyTorch implementation ready to go.

New Samplers

There have also been a couple new Stable Diffusion samplers released recently:

Generative Audio

Generative Audio

Still not enough generative AI resources for you? You’re in luck. There have also been developments in audio-based generative AI:

  • Nendo is a tool suite for producing AI-based audio applications.
  • If you need to transcribe audio to text in a hurry, Insanely Fast Whisper has got you covered.

Machine Learning and Scientific Computing

Machine Learning and Scientific Computing

Generative AI had some pretty huge releases recently. But that’s not the only corner of AI that’s been seeing progress.

Machine Learning Frameworks

Just in the last couple of weeks, we’ve had three new machine learning frameworks land on our doorstep:

  • BackboneLearn uses mixed-integer optimization to deliver high-quality machine learning models through powerful optimization algorithms (if you’re especially interested in this area, check out this book by optimization rockstar Dimitris Bertsimas)
  • SparseML is all about, well, sparse machine learning. Whether it’s for model regularization or faster inference, sparsity is generally a good thing in machine learning solutions. Great to see an entire library dedicated to sparse learning.
  • Even reinforcement learning is getting in on the action. Stable-Baselines3 – Contrib may not have the catchiest name, but they add cutting-edge RL algorithms that are compatible with the popular Stable-Baselines3 library for PyTorch.

Transformers and NLP

We’ve also seen some new libraries pop up in the world of transformers and NLP:

  • Adapters is an extension to the popular HuggingFace Transformers library that allows users to perform parameter-efficient and modular fine-tuning through the adapters hosted on AdapterHub.
  • If you’re looking to do topic modeling, definitely give BERTopic a look-see. Based on the technique outlined in this paper, BERTopic combines BERT with c-TF-IDF to create flexible clusters that can be applied to a wide variety of techniques and use cases.

Large Language Models (LLMs) and Retrieval Augmented Generation (RAG)

Large Language Models (LLMs) and Retrieval Augmented Generation (RAG)

Building RAGs and LLM Apps

Let’s start with the obvious suspects. LlamaIndex and LangChain can’t go for more than a week without adding new features. In particular, LlamaIndex v0.9 became official. Their recent additions include:

  • Fuzzy citation
  • Multi-rag
  • Datasets
  • Ingestion Pipelines
  • Transformation caching
  • Custom transformations
  • New interface for parsing and splitting text
  • Improved token counting

In contrast, LangChain has been relatively quiet lately. Their main addition after their slew of implementations around GPT-compatible agents after OpenAI’s Dev Summit is a new template for Skeleton-of-Thought.

Elsewhere in LLM Land, we have more releases:

  • For you RAG-builders out there, EMBD is a new GPU-accelerated library for computing text embeddings with the twist that it’s designed to be run entirely client-side (which could make your RAG architecture much simpler).
  • Want to get your text embeddings the traditional way, but need more speed? In that case, check out Text Embeddings Inference. It’s a HuggingFace project, so you know you’re getting good quality. It’s also written in Rust, so top-notch performance is all-but guaranteed.
  • After all the drama with ChatGPT and even the GPT API going down completely (and the recent reports of diminished GPT quality), people are thinking about backups for their LLM-powered solutions. OpenRouter is designed to do just that, routing your LLM API requests to the best available provider at the best available price.
  • Looking for a lightweight vector database for your RAG? Check out VLite V2. It’s an ultra-fast, feature-rich vector database, crafted with just NumPy in under 250 lines of Python. Long-term maintenance is up in the air, but it’s an intriguing project worthy of attention, especially for its code efficiency.
  • For faster LoRA or QLoRA fine-tuning, consider unsloth, which offers 80% faster speeds and cuts memory usage by 50%, leveraging a custom autograd engine and Triton kernels for efficient computation without sacrificing accuracy.

LLM Apps

We’ll round out this post with a couple of LLM-based apps that happened to catch our attention:

  • Self-Operating Computer Framework may look similar to Open Interpreter, but the concept of an LLM-driven computing experience is so cool that we couldn’t let the opportunity pass to give it a mention.
  • The other app that caught our eye was CrewAI, a framework for autonomous role-playing agents. This idea also exists in AI Town, but once again, we couldn’t resist mentioning something as cool as LLM-driven characters in a video game.

Signing Off

As you can see, it’s been a busy couple of weeks in AI. For readability purposes, we had to make some tough choices about apps and libraries to cut from our coverage. Our apologies to the hard-working creators out there that we didn’t get around to mentioning.

Thanks for reading! We hope you found one or two gems to tinker with. We’ll see you next time with more AI news, tutorials and perspectives.

Secret Link