Long video generation remains limited by weak controllability, spatial drift, and memory loss when objects leave and re-enter the view. We propose VoxelMem, a framework that uses 4D voxels as both controllable primitives and memory carriers. Each voxel stores time-dependent geometry, visibility, segmentation identity, and local appearance, enabling multi-granularity control from object-level layout to part-level motion. To maintain long-term consistency, VoxelMem combines local appearance propagated with voxel motion and global history frames retrieved by future voxel coverage, providing both fine-grained spatial correspondence and holistic scene references. We further extend VoxelMem to streaming generation by distilling the bidirectional model into a causal generator. Finally, we construct 4D video-voxel training data from synthetic rendered scenes and real-world videos to support multi-granularity control and memory learning.