AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What’s Needed To Teach AI Watercolour Painting Via TRL And OpenEnv? on ThorstenMeyerAI.com

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

TL;DR

An engineer has created a complete, open reproduction of Surya Narreddi’s viral watercolour painting AI model, utilizing TRL and OpenEnv. All datasets, scripts, and trained models are now publicly available, enabling further research into aesthetic reinforcement learning.

An independent engineer has successfully recreated Surya Narreddi’s viral watercolour painting language model using open-source tools TRL and OpenEnv, with all artifacts released publicly on Hugging Face. This process is detailed in the original analysis. This reproduction aims to demonstrate how reinforcement learning can optimize models based on aesthetic taste rather than verifiable answers, marking a significant step toward open AI art research.

The reproduction, built with the TRL framework and OpenEnv environment, retrains a Qwen language model with reinforcement learning against an aesthetic reward mix, illustrating how AI models can be trained for creative tasks. The project includes the release of datasets, environment scripts, training code, and trained models under an open license, making it accessible for further experimentation and development.

It follows the original project by Surya Narreddi, whose August 23 viral video showcased AI-generated watercolour paintings that drew over 1.5 million views. The original project trained a language model to generate JavaScript code that produces watercolour-style images, with style enforced through strict method restrictions. The reproduction replicates this process, using a reward system that combines compilation checks, code length, style judgment, and a preference model trained on human choices.

The core innovation lies in applying reinforcement learning to optimize aesthetic qualities, rather than solving problems with clear-cut answers, as discussed in this detailed analysis. The project tests whether models can learn to produce visually pleasing art based on human-like preferences, rather than objective correctness, a question that has gained interest amid the rise of AI art tools.

At a glance
updateWhen: published March 2024
The developmentAn independent engineer has reproduced Surya Narreddi’s watercolour painting AI model, releasing all code, datasets, and trained models on Hugging Face, following the original viral video.
At a glance
reportWhen: published after the 23 August viral vid…
The developmentA fully open reproduction of Surya Narreddi’s viral watercolour-painting coding model — including the RL environment, reference dataset, training scripts and trained models — has been published, built with TRL and OpenEnv and running entirely on Hugging Face infrastructure.

Implications for AI-Generated Art and Aesthetic Optimization

This open reproduction advances transparency in AI art generation by providing accessible tools and datasets, enabling researchers to experiment with reinforcement learning based on aesthetic preferences. It also demonstrates that open-source frameworks like TRL and OpenEnv can facilitate complex training pipelines, potentially accelerating innovation in AI-driven creativity.

By focusing on aesthetic taste rather than verifiable correctness, this work challenges traditional reinforcement learning paradigms and opens new avenues for exploring subjective quality in AI outputs. The project’s open artifacts lower barriers for researchers and artists interested in developing models that learn from human preferences, potentially influencing future AI art tools and interactive systems.

Amazon

AI art creation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background and Evolution of AI Art and Reinforcement Learning Approaches

The project sits within a lineage of early AI art experiments, from DeepDream (2015) to neural network portraits by Mario Klingemann, emphasizing exploration of generative algorithms as artistic tools. The original watercolour model by Surya Narreddi gained viral attention in August 2023, showcasing AI-generated paintings that mimic traditional watercolour techniques using code-driven brushstrokes.

Prior efforts in AI art training often relied on curated datasets, human annotations, or fixed style transfer methods. Narreddi’s approach introduced reinforcement learning based on human preferences, aiming to produce more natural and stylistically pleasing artworks. The open reproduction builds on this by providing full transparency and reproducibility, which were lacking in the initial project.

This development aligns with broader trends in AI research, where open models and datasets are increasingly prioritized to foster collaborative progress and avoid proprietary bottlenecks. The project also echoes earlier artistic experiments with curated datasets, such as Anna Ridler’s tulip dataset, but extends into dynamic reinforcement learning tailored for aesthetic judgment.

“By fully open-sourcing the datasets, environment scripts, and trained models, we aim to democratize research into aesthetic reinforcement learning and AI art generation.”

— Thorsten Meyer, project developer

Amazon

watercolor painting AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Model Fidelity and Artistic Quality

It remains unclear how closely the reproduction’s outputs match the original viral paintings in terms of artistic quality and style fidelity. The project’s authors have not provided quantitative comparisons or detailed evaluations of the generated artworks against the original.

Additionally, the effectiveness of the different reward mixes in optimizing aesthetic taste remains unconfirmed, as no definitive results or user studies are yet available. The full technical report from Narreddi, which might clarify these points, has not been published.

It is also uncertain whether this open approach will scale to more complex artistic styles or broader visual domains, as the current focus is on watercolour-like images and JavaScript code outputs.

Amazon

AI art training datasets

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Research and Community Engagement Opportunities

The immediate next step is the release of Narreddi’s full technical report, which is expected to detail the training process, reward design, and evaluation metrics. This will help clarify how well the reproduction aligns with the original project and its artistic outputs.

Researchers and artists are encouraged to experiment with the open artifacts, testing different reward configurations, datasets, and model architectures to explore the potential of aesthetic reinforcement learning further. Collaborative efforts could extend the approach to other artistic styles or interactive applications.

Moreover, ongoing development of open frameworks like TRL and OpenEnv will likely facilitate more complex and nuanced models capable of learning subjective qualities, fostering a broader community around AI-driven creative tools.

Amazon

reinforcement learning art models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main achievement of this open reproduction?

The project provides a complete, open-source pipeline for training a watercolour-style AI model using reinforcement learning based on human-like aesthetic preferences, including datasets, scripts, and trained models.

How does reinforcement learning differ in this project compared to traditional AI training?

Instead of optimizing for verifiable answers or fixed metrics, the model learns from human-like aesthetic preferences, aiming to produce more stylistically pleasing artworks based on subjective taste.

Can this approach be applied to other artistic styles?

Potentially yes, but further experimentation is needed to assess how well the reinforcement learning framework generalizes beyond watercolour paintings and JavaScript code outputs.

What are the limitations of the current reproduction?

It is not yet clear how closely the generated artworks match the original viral paintings in style or quality, and the effectiveness of different reward mixes has not been quantitatively evaluated.

When will the full technical report be available?

According to the project’s source, the full report is expected soon, which will provide detailed insights into the training process and evaluation results.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Jack Clark Says It Out Loud — Reading the Co-Founder’s 60%/2028 Estimate on Automated AI R&D

Anthropic’s co-founder Jack Clark states there is a 60%+ probability that AI systems capable of self-creating successors will emerge by 2028, marking a significant policy forecast.

Uber to open 2 campuses in India to support product development, operations

Uber plans to open two campuses in Bengaluru and Hyderabad by 2027, boosting its engineering and AI capabilities in India, and partnering with Adani for a data center.

GenCAD

GenCAD is a new AI model that generates parametric 3D CAD models and their command histories from images, advancing automated design capabilities.

What’s Behind Sony’s Lawsuit Against Anthropic Over AI-Generated Music?

Sony accuses Anthropic of using Sony music without permission in training its AI model Claude, seeking up to $150,000 per song. Details remain unclear.