Twilight’s Flash Fiction brings together more
than six hundred short stories exploring science fiction, fantasy, memory,
humour and philosophical speculation. The aim of the audiovisual project was to
transform some of them into short videos that could be published on YouTube and
encourage viewers to read the original stories.
To achieve this, I
combined generative artificial intelligence tools with conventional editing
applications. The process evolved as new creative requirements—and new
technical problems—appeared.
Tools Used
ChatGPT was used to analyse the
stories, identify the main scenes, condense the narration, prepare visual
scripts, create prompts, and review the Spanish and English versions.
The first images were
generated with Stable Diffusion through Hugging Face. Later, I moved the
process to my local computer using ComfyUI and SDXL Base 1.0, running on
an HP OMEN Max 16 equipped with an NVIDIA GeForce RTX 5070 Ti Laptop GPU with
12 GB of memory.
The final editing was
completed in Canva, where I added the images, text, transitions, camera
movements and music, and adjusted the duration of each scene. The finished
videos were published on YouTube and subsequently embedded in the blog through
Blogger.
Workflow
1. Selecting and Adapting the Story
The first step was to
identify the core of the story: the image, conflict or revelation that had to
survive the adaptation. We then divided the text into a small number of scenes.
Trying to illustrate
every sentence produced a sequence of slides rather than an audiovisual
narrative. The result improved when we selected only a few meaningful moments
and allowed viewers to fill in the transitions themselves.
2. Preparing the Prompts
Each scene was
transformed into a visual description that included:
- Characters and clothing.
- Location and historical period.
- Lighting and colour palette.
- Camera position and distance.
- Aspect ratio and
general style.
- Elements that should
not appear.
Maintaining a common
structure across the prompts helped reduce visual differences between the
generated images.
3. Local Generation with ComfyUI
The basic ComfyUI
workflow consisted of the following nodes:
Load
Checkpoint → CLIP Text Encode → Empty Latent Image → KSampler → VAE Decode →
Save Image
Both positive and
negative prompts were used, and the resulting images were saved as PNG files.
For horizontal videos, I worked with resolutions such as 1344 × 768 before
preparing the final export at 1920 × 1080.
Saving the TXT2IMG_BASE and IMG2IMG_BASE
workflows prevented me from having to rebuild the process after each session or
computer restart.
4. Editing in Canva
The selected images were
imported into Canva, where I adjusted:
- Scene duration.
- Cropping and framing.
- Slow zoom effects.
- Transitions.
- Titles and captions.
- Music.
- MP4 export settings.
Titles were always added
in Canva. Asking the image generator to produce text usually resulted in
distorted letters or nonexistent words.
Problems
Encountered
|
Problem
|
Solution
|
|
Changes in faces,
clothing or visual style
|
Reuse descriptions and
maintain a common visual guide
|
|
Scenes with incompatible colours
|
Define a palette in
advance, such as midnight blue
|
|
Black images during VAE
decoding
|
Set the batch size to 1
and launch ComfyUI with --fp32-vae.
|
|
Defective text inside
generated images
|
Generate images without
text and add it later in Canva
|
|
CUDA illegal instruction in KSampler
|
Work in small batches
and save every valid result
|
|
Computer restarts after
several generations
|
Preserve the workflows
and export approved images immediately
|
|
Unwanted cropping after
changing formats
|
Choose 16:9 or 9:16
before starting image generation
|
Lessons Learned
The main conclusion is
that the story must govern the technology. A spectacular image is of
little value if it breaks the visual continuity of the video.
I also learned that a
reusable workflow is more valuable than a lucky prompt, that a few coherent
images work better than many unrelated ones, and that the final aspect ratio
must be chosen before image generation begins.
Artificial intelligence
allows a writer working alone to produce a small audiovisual narrative.
However, it does not remove authorship or creative decision-making. On the
contrary, the more powerful the tools become, the more important it is to know
exactly which story we want to tell.
Cómo
convertí relatos de Twilight’s
Flash Fiction en vídeos
con IA
Herramientas, flujo de trabajo y
problemas reales
Twilight’s Flash Fiction reúne más de
seiscientos relatos breves de ciencia ficción, fantasía, memoria, humor y
especulación filosófica. El propósito del proyecto audiovisual era convertir
algunos de ellos en vídeos breves que pudieran publicarse en YouTube y conducir
al espectador hacia el texto original.
Para hacerlo combiné herramientas de inteligencia
artificial generativa con aplicaciones convencionales de edición. El proceso
fue evolucionando a medida que aparecían nuevas necesidades y, también, nuevos
problemas.
Herramientas
utilizadas
ChatGPT se utilizó
para analizar los relatos, identificar las escenas principales, condensar la
narración, preparar los guiones visuales, elaborar los prompts y revisar las
versiones española e inglesa.
Las primeras imágenes se generaron con Stable
Diffusion a través de Hugging Face. Más adelante trasladé el proceso al
ordenador local mediante ComfyUI y SDXL Base 1.0, ejecutados en un HP
OMEN Max 16 con una tarjeta NVIDIA RTX 5070 Ti Laptop de 12 GB.
El montaje se realizó en Canva, donde
incorporé las imágenes, los textos, las transiciones, los movimientos de
cámara, la música y la duración de cada escena. Los vídeos terminados se
publicaron en YouTube y se incorporaron posteriormente al blog mediante
Blogger.
Flujo de
trabajo
1. Selección y
adaptación del relato
El primer paso consistía en localizar el núcleo
del cuento: la imagen, el conflicto o la revelación que debía conservarse.
Después dividíamos el relato en un pequeño número de escenas.
Intentar ilustrar cada frase producía una
sucesión de diapositivas, no una narración audiovisual. El resultado mejoraba
al seleccionar pocos momentos significativos y permitir que el espectador
completara las transiciones.
2. Preparación
de los prompts
Cada escena se transformaba en una descripción
visual que incluía:
- Personajes y vestuario.
- Lugar y época.
- Iluminación y paleta cromática.
- Distancia y posición de la cámara.
- Formato y estilo general.
- Elementos que no debían aparecer.
Mantener una estructura común en los prompts
ayudaba a reducir las diferencias entre las imágenes.
3. Generación
local con ComfyUI
El flujo básico utilizado en ComfyUI estaba
formado por los siguientes nodos:
Checkpoint
Loader → CLIP Text Encode → Empty Latent Image → KSampler → VAE Decode → Save
Image
Se empleaban prompts positivos y negativos, y las
imágenes se guardaban en PNG. Para los vídeos horizontales se trabajó con
proporciones como 1344 × 768, preparando posteriormente la exportación final a
1920 × 1080.
Guardar los flujos TXT2IMG_BASE e IMG2IMG_BASE evitaba tener
que reconstruir el proceso después de cada sesión o reinicio.
4. Montaje en
Canva
Las imágenes seleccionadas se importaban en
Canva. Allí se ajustaban:
- Duración de las escenas.
- Recortes y encuadres.
- Acercamientos lentos.
- Transiciones.
- Títulos y textos.
- Música.
- Exportación a MP4.
Los títulos se añadían siempre en Canva. Pedir al
generador de imágenes que escribiera texto producía letras deformadas o
palabras inexistentes.
Problemas
encontrados
|
Problema
|
Solución aplicada
|
|
Cambios de rostro, vestuario o estilo
|
Repetir descripciones y mantener una guía visual común
|
|
Escenas con colores incompatibles
|
Definir previamente una paleta, como el azul noche
|
|
Imágenes negras durante la decodificación VAE
|
Establecer el tamaño del lote en 1 e iniciar ComfyUI con --fp32-vae.
|
|
Texto defectuoso dentro de las imágenes
|
Generar sin letras y añadirlas después en Canva
|
|
Error CUDA illegal instruction in KSampler
|
Trabajar en lotes pequeños y guardar cada resultado válido
|
|
Reinicios después de varias generaciones
|
Conservar los flujos y exportar inmediatamente las imágenes
|
|
Recortes al cambiar el formato
|
Elegir 16:9 o 9:16 antes de comenzar la generación
|
Lecciones
aprendidas
La principal conclusión es que el relato debe
gobernar la tecnología. Una imagen espectacular no resulta útil si rompe la
continuidad del vídeo.
También aprendí que un flujo de trabajo
reutilizable vale más que un prompt afortunado, que unas pocas imágenes
coherentes funcionan mejor que muchas imágenes inconexas y que el formato final
debe decidirse antes de generar.
La inteligencia artificial permite que un
escritor que trabaja solo pueda producir una pequeña narración audiovisual. Sin
embargo, no elimina la autoría ni las decisiones creativas. Al contrario:
cuanto más potentes son las herramientas, más importante resulta saber qué
historia se quiere contar.