# 2023-11-12 13:17:35 Try IRIs but with pretrained transformer with LoRA adapter - [x] first can I run it yes with a 1/2 batch size - [ ] then can I add 3B with adapter... ```sh poetry install . ./.venv/bin/activate python src/main.py env.train.id=BreakoutNoFrameskip-v4 common.device=cuda:0 wandb.mode=offline # or for quick debug WANDB_MODE=disabled python -m pdb src/main.py env.train.id=BreakoutNoFrameskip-v4 ``` ```sh # TODO use this code to load a transformer, and other code from my bigvae repo https://github.com/wassname/bigvae_wm def load_model(config, device='cuda'): tokenizer = AutoTokenizer.from_pretrained(config.model_name, trust_remote_code=True) tokenizer.padding_side = "left" if tokenizer.pad_token is None: tokenizer.pad_token = tokenizer.eos_token bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_quant_type="nf4", bnb_4bit_use_double_quant=True, ) base_model = AutoModelForCausalLM.from_pretrained( config.model_name, device_map={"": device}, quantization_config=bnb_config, torch_dtype=torch.bfloat16, trust_remote_code=True ) peft_config = peft.LoraConfig( peft.TaskType.CAUSAL_LM, inference_mode=False, r=config.rank, lora_alpha=8, lora_dropout=config.dropout, target_modules=[ "self_attn.q_proj", "self_attn.k_proj", "self_attn.v_proj", "self_attn.o_proj", "mlp.gate_proj", "mlp.up_proj", "mlp.down_proj", ], ) base_model_peft = peft.get_peft_model(base_model, peft_config) vae_model = BigVAE( base_model_peft, device, peft_config, z_dim=config.z_dim, ) if config.start_from: vae_model.load_pretrained(config.start_from) base_model_peft.requires_grad_(False) vae_model.vae_head.requires_grad_(False) vae_model.vae_head.w_d.requires_grad_() router = BigVAERouter(base_model_peft, vae_model, device, peft_config) if config.start_from: router.load_pretrained(config.start_from, is_trainable=True) print(router.model.print_trainable_parameters()) router.model.set_adapter("router") ``` Debugging: batch['observations'].shape torch.Size([16, 20, 3, 64, 64]) obs_tokens.shape torch.Size([16, 20, 16]) https://vscode.dev/github/wassname/iris_bigvae/blob/just_llms2/src/models/world_model.py#L105 tokens tensor([[222, 222, 222, ..., 409, 55, 2], [222, 222, 222, ..., 409, 139, 1], [222, 222, 222, ..., 168, 190, 3], ..., [222, 222, 222, ..., 168, 55, 0], [222, 222, 222, ..., 237, 190, 3], [222, 222, 222, ..., 168, 55, 0]], device='cuda:0') tokens.shape torch.Size([16, 340]) where 16 is the batch size. 340 is the step size?. actions was 16,20 int tokens.shape int torch.Size([16, 340]) sequences.shape float32 torch.Size([16, 340, 256]) transfrmer x.shape torch.Size([16, 340, 256]) # 2023-11-12 16:58:37 So I got it training, but during imagination it passes in a single token with no past steps. But the slicer seems to need at least on block? And so I get none? hmm it's because num_kept_tokens is 16 not 1. So there should be a whole block passed in ? wait apparently it's also a problem in the normal repo.... I confuse! maybe it's my config! maybe I need >larger than block size. nope hmm it still happens in the original repo with my debug params. maybe it's my debug params ... trying a full run without my debug params... note trains.world_model.batch_num_samples:4 fill 20GB gpu ram for the 3b stability ai llm ok even with a full run I get the error. I think it's a bug in the original repo. I'll try to debug it there. Epoch 51 / 600 Experience collection (train_dataset): 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 200/200 [00:03<00:00, 59.91it/s] Training tokenizer: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 200/200 [00:17<00:00, 11.53it/s] Training world_model: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 200/200 [02:11<00:00, 1.53it/s] Training actor_critic: 0%| | 0/200 [00:00 (B, K) ``` - Embed $Emb(t_0) = z_0$ ```py embedded_tokens = self.tokenizer.embedding(self.obs_tokens) # (B, K, E) z = rearrange(embedded_tokens, 'b (h w) e -> b e h w', h=int(np.sqrt(self.num_observations_tokens))) ``` - Dynamics $D(z_0, a_0) = z_1$ ```py outputs_wm = self.world_model(token, past_keys_values=self.keys_values_wm) ``` - Decoder $D(z_0, a_0) = x_1$ ```py rec = self.tokenizer.decode(z, should_postprocess=True) # (B, C, H, W) ``` but we have tokens vs z Questions: - wait why are we just passing in "action_token" to the transformer and not obs? that must have obs in it right... right??? confirm - in iris-delta how did they pass everything in? I guess obs_prev was tokenized too? I think the slices are annoying so maybe I should just pass things seperatly # 2023-11-24 10:56:40 If I unfreeze the whole transformer, it seem to learn the most obvious dynamics (the next latent space is the same as the last). To summarize - with Qlora it didn't learn that - with unfrozen head it didn't - when training transformer and obs embedding together it did not (frozen llm embeddings) no it didn't work with tokenizer sep hmm