AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What’s Needed To Teach AI Watercolour Painting Via TRL And OpenEnv? on ThorstenMeyerAI.com

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

TL;DR

An engineer has created a complete, open reproduction of Surya Narreddi’s viral watercolour painting AI model, utilizing TRL and OpenEnv. All datasets, scripts, and trained models are now publicly available, enabling further research into aesthetic reinforcement learning.

An independent engineer has successfully recreated Surya Narreddi’s viral watercolour painting language model using open-source tools TRL and OpenEnv, with all artifacts released publicly on Hugging Face. This process is detailed in the original analysis. This reproduction aims to demonstrate how reinforcement learning can optimize models based on aesthetic taste rather than verifiable answers, marking a significant step toward open AI art research.

The reproduction, built with the TRL framework and OpenEnv environment, retrains a Qwen language model with reinforcement learning against an aesthetic reward mix, illustrating how AI models can be trained for creative tasks. The project includes the release of datasets, environment scripts, training code, and trained models under an open license, making it accessible for further experimentation and development.

It follows the original project by Surya Narreddi, whose August 23 viral video showcased AI-generated watercolour paintings that drew over 1.5 million views. The original project trained a language model to generate JavaScript code that produces watercolour-style images, with style enforced through strict method restrictions. The reproduction replicates this process, using a reward system that combines compilation checks, code length, style judgment, and a preference model trained on human choices.

The core innovation lies in applying reinforcement learning to optimize aesthetic qualities, rather than solving problems with clear-cut answers, as discussed in this detailed analysis. The project tests whether models can learn to produce visually pleasing art based on human-like preferences, rather than objective correctness, a question that has gained interest amid the rise of AI art tools.

At a glance
updateWhen: published March 2024
The developmentAn independent engineer has reproduced Surya Narreddi’s watercolour painting AI model, releasing all code, datasets, and trained models on Hugging Face, following the original viral video.
At a glance
reportWhen: published after the 23 August viral vid…
The developmentA fully open reproduction of Surya Narreddi’s viral watercolour-painting coding model — including the RL environment, reference dataset, training scripts and trained models — has been published, built with TRL and OpenEnv and running entirely on Hugging Face infrastructure.

Implications for AI-Generated Art and Aesthetic Optimization

This open reproduction advances transparency in AI art generation by providing accessible tools and datasets, enabling researchers to experiment with reinforcement learning based on aesthetic preferences. It also demonstrates that open-source frameworks like TRL and OpenEnv can facilitate complex training pipelines, potentially accelerating innovation in AI-driven creativity.

By focusing on aesthetic taste rather than verifiable correctness, this work challenges traditional reinforcement learning paradigms and opens new avenues for exploring subjective quality in AI outputs. The project’s open artifacts lower barriers for researchers and artists interested in developing models that learn from human preferences, potentially influencing future AI art tools and interactive systems.

Amazon

AI art creation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background and Evolution of AI Art and Reinforcement Learning Approaches

The project sits within a lineage of early AI art experiments, from DeepDream (2015) to neural network portraits by Mario Klingemann, emphasizing exploration of generative algorithms as artistic tools. The original watercolour model by Surya Narreddi gained viral attention in August 2023, showcasing AI-generated paintings that mimic traditional watercolour techniques using code-driven brushstrokes.

Prior efforts in AI art training often relied on curated datasets, human annotations, or fixed style transfer methods. Narreddi’s approach introduced reinforcement learning based on human preferences, aiming to produce more natural and stylistically pleasing artworks. The open reproduction builds on this by providing full transparency and reproducibility, which were lacking in the initial project.

This development aligns with broader trends in AI research, where open models and datasets are increasingly prioritized to foster collaborative progress and avoid proprietary bottlenecks. The project also echoes earlier artistic experiments with curated datasets, such as Anna Ridler’s tulip dataset, but extends into dynamic reinforcement learning tailored for aesthetic judgment.

“By fully open-sourcing the datasets, environment scripts, and trained models, we aim to democratize research into aesthetic reinforcement learning and AI art generation.”

— Thorsten Meyer, project developer

Amazon

watercolor painting AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Model Fidelity and Artistic Quality

It remains unclear how closely the reproduction’s outputs match the original viral paintings in terms of artistic quality and style fidelity. The project’s authors have not provided quantitative comparisons or detailed evaluations of the generated artworks against the original.

Additionally, the effectiveness of the different reward mixes in optimizing aesthetic taste remains unconfirmed, as no definitive results or user studies are yet available. The full technical report from Narreddi, which might clarify these points, has not been published.

It is also uncertain whether this open approach will scale to more complex artistic styles or broader visual domains, as the current focus is on watercolour-like images and JavaScript code outputs.

Amazon

AI art training datasets

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Research and Community Engagement Opportunities

The immediate next step is the release of Narreddi’s full technical report, which is expected to detail the training process, reward design, and evaluation metrics. This will help clarify how well the reproduction aligns with the original project and its artistic outputs.

Researchers and artists are encouraged to experiment with the open artifacts, testing different reward configurations, datasets, and model architectures to explore the potential of aesthetic reinforcement learning further. Collaborative efforts could extend the approach to other artistic styles or interactive applications.

Moreover, ongoing development of open frameworks like TRL and OpenEnv will likely facilitate more complex and nuanced models capable of learning subjective qualities, fostering a broader community around AI-driven creative tools.

Amazon

reinforcement learning art models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main achievement of this open reproduction?

The project provides a complete, open-source pipeline for training a watercolour-style AI model using reinforcement learning based on human-like aesthetic preferences, including datasets, scripts, and trained models.

How does reinforcement learning differ in this project compared to traditional AI training?

Instead of optimizing for verifiable answers or fixed metrics, the model learns from human-like aesthetic preferences, aiming to produce more stylistically pleasing artworks based on subjective taste.

Can this approach be applied to other artistic styles?

Potentially yes, but further experimentation is needed to assess how well the reinforcement learning framework generalizes beyond watercolour paintings and JavaScript code outputs.

What are the limitations of the current reproduction?

It is not yet clear how closely the generated artworks match the original viral paintings in style or quality, and the effectiveness of different reward mixes has not been quantitatively evaluated.

When will the full technical report be available?

According to the project’s source, the full report is expected soon, which will provide detailed insights into the training process and evaluation results.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Sony ups its new A7R VI to 66.8 megapixels and jumps the price to $4,500

Sony announces the A7R VI with a 66.8-megapixel stacked sensor, new features, and a $4,500 price, launching in June. What this means for photographers.

Trends in High‑Speed Inkjet Technology

Breaking advancements in high-speed inkjet technology are transforming industrial printing—discover how these innovations could change your printing future.

Waymo suspends freeway driving amid safety concerns

Waymo has halted freeway driving in all US markets due to safety issues related to construction zones and recent flooding incidents, impacting high-speed autonomous trips.

Can ByteDance’s $29.6B Loan Transform AI Industry On A Worldwide Scale?

ByteDance reports securing a $29.6 billion loan to fund its worldwide AI expansion, though key details about the loan’s terms and use remain undisclosed.