Your browser cannot display this tag cloud.

Visual Portfolio, Posts & Image Gallery for WordPress

Warpfusion: Smooth animations

A technique for maintaining a consistent style and reducing flicker in AI-generated animation.*Github https://github.com/Sxela/DiscoDiffusion-Warp*Colab https://colab.research.google.com/drive/15D2WIF_vE2l48ddxEx45cM3RykZwQXM8?usp=sharing

Time-Travel Rephotography: Portrait restoration

A method for restoring old portraits that goes beyond colourisation to restructure the entire photograph as though it had been taken with a modern camera. It essentially maps the image into the latent space of a StyleGAN2 trained on modern high-resolution portraits, achieving denoising, colourisation and super-resolution simultaneously.*Web https://time-travel-rephotography.github.io/*Colab https://colab.research.google.com/drive/15D2WIF_vE2l48ddxEx45cM3RykZwQXM8?usp=sharing*Paper https://arxiv.org/pdf/2012.12261.pdf

SRT – Scene Representation Transformer: Full 3D scene synthesis

From a small number of images, this system reconstructs a scene in real time, “hallucinating” the parts it does not know to make the scene coherent.*Web https://srt-paper.github.io/*Colab https://colab.research.google.com/github/srt-paper/srt-paper.github.io/blob/main/multi_shapenet.ipynb*Github https://github.com/stelzner/srt*Paper https://arxiv.org/abs/2111.13152

Dream Fields: Text to 3D

Dreamfields is a system that creates 3D models from text prompts. *Colab: https://colab.research.google.com/drive/1TjCWS2_Q0HJKdi9wA2OSY7avmFUQYGje?usp=sharing*Web https://ajayj.com/dreamfields*Github https://github.com/google-research/google-research/tree/master/dreamfields*Paper https://arxiv.org/abs/2112.01455

VGPNN: Video generation and manipulation

A tool for manipulating video — temporal extension, summarisation, resizing and style transfer, among other tasks — without deep learning. Instead, it uses a more classical approach based onGranot et al. (2021) enabling substantially faster processing. * Web https://nivha.github.io/vgpnn/ *github https://github.com/nivha/single_video_generation *paper https://arxiv.org/abs/2109.08591

Spleeter4Max: A tool for separating audio tracks in Ableton using Spleeter

Spleeter4Max is a tool that allows for the separation of audio tracks in Ableton using Spleeter, making the process of breaking down songs into different elements such as vocals and instruments easier.*Ableton https://github.com/diracdeltas/spleeter4max/releases *Max https://github.com/diracdeltas/spleeter4max/tree/feature/native-spleeter#spleeter-for-max-native-version *Tutorial https://www.youtube.com/watch?v=4pcJoI5CUOA

InstColorization: Image colourisation

Instance-aware image colourisation: in other words, an AI system that colours black-and-white images after identifying the objects they contain, working on each part separately. *Web https://wandb.ai/wandb/instacolorization/reports/Overview-Instance-Aware-Image-Colorization---VmlldzoyOTk3MDI *Colab https://colab.research.google.com/github/ericsujw/InstColorization/blob/master/InstColorization.ipynb *Github https://github.com/ericsujw/InstColorization *Paper https://cgv.cs.nthu.edu.tw/InstColorization_data/InstaColorization.pdf

Deezer’s Spleeter: A Source Separation Library with Pretrained Models

Spleeter is a source separation library developed by Deezer that includes pretrained models. Written in Python and utilizing TensorFlow, it facilitates the separation of audio tracks into different components such as vocals, drums, bass, among others, with high processing speed, especially on a GPU. It can be used straight from the command line or integrated into development pipelines as a Python library.*Web https://pypi.org/project/spleeter/ *Colab https://colab.research.google.com/github/deezer/spleeter/blob/master/spleeter.ipynb *Manual https://github.com/deezer/spleeter/wiki *Github https://github.com/deezer/spleeter

MuseNet: A deep neural network generating musical compositions with various instruments and styles.

MuseNet is a deep neural network developed by OpenAI that can generate 4-minute musical compositions with up to 10 different instruments, blending styles from country to Mozart to The Beatles. MuseNet was not explicitly programmed with an understanding of music, but instead discovered patterns of harmony, rhythm, and style by learning to predict the next token in hundreds of thousands of MIDI files. It utilizes a general-purpose unsupervised technology similar to GPT-2. *Web https://openai.com/research/musenet *Colab https://colab.research.google.com/github/asigalov61/OpenAI-MuseNet-Colab-Notebook/blob/main/OpenAI_MuseNet_Colab_Notebook.ipynb

MediaPipe Hands: Real-time, cross-platform hand pose tracking

MediaPipe Hands is a high-fidelity hand and finger tracking solution. It uses machine learning (ML) to infer 21 3D hand landmarks from a single frame, achieving real-time performance on a mobile phone and even on the web for multiple hands. * Web https://google.github.io/mediapipe/solutions/hands *Demo https://rdtr01.xl.digital/ *Demo + Code https://codepen.io/mediapipe/pen/RwGWYJw *Github https://github.com/google/mediapipe

WaveNet: A Generative Model for Raw Audio

WaveNet is a deep generative model capable of creating raw audio waveforms, able to generate speech that mimics any human voice, sounding more natural than existing text-to-speech systems. Moreover, it can synthesize other types of audio signals, such as music, automatically generating piano pieces with striking results.*Web https://www.deepmind.com/research/highlighted-research/wavenet *Demo https://play.ht/text-to-speech-voices/google-wavenet/ *Colab https://colab.research.google.com/github/olaviinha/WaveNet/blob/master/WaveNet.ipynb#scrollTo=Zidq9-vvLF7l *Github https://github.com/ibab/tensorflow-wavenet